• cman6@lemmy.world
    link
    fedilink
    English
    arrow-up
    19
    ·
    7 days ago

    😂 Some of this sounds like a child wrote it (emphasis mine):

    USHERING IN SUPER INTELLIGENCE: … the executive branch to recognize the continuously advancing technological frontier and the limitless promise it offers the American people

    RECOGNIZING THE IMMENSE OPPORTUNITY OF SI: The United States is the world leader in Super Intelligence…

    Well you would be given that no one else calls it that.

    But also, have you heard of Qwen? Last time I checked it was as good if not better than American models! (I admit I could be wrong about that now as I haven’t checked benchmarks in a few months now)

    • percent@infosec.pub
      link
      fedilink
      English
      arrow-up
      5
      arrow-down
      1
      ·
      7 days ago

      It’s definitely better than American open-weight models. The gap between American flagship models and Chinese open-weight models is closing quickly though.

      With the token prices of Chinese models, I seldom reach for American models anymore.

      • FaceDeer@fedia.io
        link
        fedilink
        arrow-up
        6
        ·
        7 days ago

        I’ve spent the last week or so coding an application using nothing but Qwen3.8-27B. I was curious how much I could do entirely with a local model running on my own personal hardware, and it turns out the answer was “everything.”

        Granted, it’s not the fanciest application. But it’s fun.

        • percent@infosec.pub
          link
          fedilink
          English
          arrow-up
          1
          ·
          6 days ago

          Yep, I’ve been using that same model pretty heavily too, running on ExLlamaV3. It’s very impressive for such a little model!

          • FaceDeer@fedia.io
            link
            fedilink
            arrow-up
            1
            ·
            6 days ago

            I’m really looking forward to seeing what Qwen4 is like. 3.8 was just 3.6 with additional training, my understanding is that Qwen4 is new from the ground up.

            • percent@infosec.pub
              link
              fedilink
              English
              arrow-up
              1
              ·
              6 days ago

              I hope they make a MoE version. I get like 50 tokens/sec with Qwen3.6-35B-A3B, but can’t rely on it quite as much as Qwen3.8-27B.

              • FaceDeer@fedia.io
                link
                fedilink
                arrow-up
                1
                ·
                6 days ago

                Sadly they don’t seem to have 36B-A3B models on their roadmap any more, they didn’t do one for 3.8 either. I agree it was a nice sweet spot between speed and capability, I still use the 3.6 version of 36B-A3B for larger-scale local work. Maybe someone else will aim for that. Or Qwen4 will do something new with the architecture that makes it unnecessary.

                • percent@infosec.pub
                  link
                  fedilink
                  English
                  arrow-up
                  1
                  ·
                  6 days ago

                  My Hermes instance periodically checks the overall open-weight LLM landscape, and it recently recommended trying Ornith 1.5 35B-A3B.

                  I haven’t tried it yet, but on paper, it sounds like it has some potential.

                  • FaceDeer@fedia.io
                    link
                    fedilink
                    arrow-up
                    1
                    ·
                    6 days ago

                    Heh, I downloaded that one just recently, I read that it was good at natural prose and I’ve been working on a little pet project to make a framework for auto-writing short stories based on a simple premise. Haven’t tested it extensively yet though.