14 comments

  • kamranjon 15 minutes ago
    Dwarfstar already supports this, curious how it compares, but I use the q4 quant daily and it works really well.

    https://github.com/antirez/ds4/blob/main/docs/MODELS.md#qwen...

  • nialv7 6 minutes ago
    There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...
  • Neywiny 1 minute ago
    I don't like that some configuration is fine via arguments and others by environment variable. I've noticed LLMs like doing this. And even more, like hallucinating such things. To me the advantage of AI coding is that the boilerplate of command line arguments and passing them around becomes trivial instead of tedious.
  • deadbunny 44 minutes ago
    > Set up Strata on this PC for me: https://github.com/Niko1221/Strata - follow docs/AI_SETUP.md in that repository.

    And I thought piping to bash was bad

    • gchamonlive 31 minutes ago
      Piping to bash is definitely worse because there is no plan mode in bash. Agents also normally don't execute anything transparently, at worst you'll see it doing something weird in the logs.
    • snehesht 43 minutes ago
      Yeah, I was surprised at first then had to dig through setup.py and setup.sh files to figure out.
  • ryan_glass 12 minutes ago
    Anyone know how it compares to GLM 5.3 for real world use?
  • snehesht 1 hour ago
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    https://huggingface.co/Qwen/Qwen3.8-Flash-Next

    • roscas 14 minutes ago
      Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

      This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.

      I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

      I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

      This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.

    • thatsabadlook 32 minutes ago
      Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me
    • proc0 1 hour ago
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      • incognito124 53 minutes ago
        Qwen 3.8 flash next is way better than 27B. It's so good I dont even use claude anymore
        • mickeyp 47 minutes ago
          I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

          It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

        • snehesht 51 minutes ago
          Yeah I agree, I'm running it with Pi didn't notice much difference compared to lower tier models and the speed, of course.
          • nicce 34 minutes ago
            I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.
      • thatsabadlook 30 minutes ago
        Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.
        • geye1234 23 minutes ago
          I find 27B more accurate -- maybe because I'm running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
          • PcChip 0 minutes ago
            Spelling mistakes?

            What inference engine are you using for flag next?

  • Tepix 22 minutes ago
    Q2 quantization. Not interested.
  • prettyblocks 38 minutes ago
    I've been playing with this on a 3090 and it FLIES. Does a pretty good job too on the tasks I've thrown at it (php code base security audits).
  • gdevenyi 52 minutes ago
    I had this working with the FreeToken inference engine a month ago when they launched.

    https://github.com/FlashML-org/FreeToken

  • hypfer 37 minutes ago
    Is these another one of those repos where it turns out that claude decided to quant the KV cache to q4 or smaller?

    The Readme doesn't say, but it's all AI generated, so..

  • esafak 1 hour ago
    Has anyone calculated the effective intelligence of these quantized models?
    • nsagent 22 minutes ago
      See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].

        We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation
      
      This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.

      [1]: https://arxiv.org/abs/2608.08188

      • merbanan 15 minutes ago
        I created a pruned experts model of the q2 quant, while it gave good performance on limited hardware there was severe quality degradation.
    • mkl 53 minutes ago
      There's some info in the README, including:

      > Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.

      https://github.com/Niko1221/Strata#which-model-should-i-pick

      • nicce 37 minutes ago
        I wonder how this Coder compares to Qwen 3.8 27B. Can it be really better since they are competitive for same memory requirements?
      • nisarg2 9 minutes ago
        92% is halfway to 99%

        Holds up pretty well

      • javier2 49 minutes ago
        ok that is getting interesting!
  • panny 28 minutes ago
    I'm far less interested in how good a big expensive model is on hardware 99% of people can't afford and would rather see what runs best on a chromebook or mobile phone with 8GB of RAM.
    • somenameforme 5 minutes ago
      The card in question here had an initial MSRP of $1600. It's been bumped up by the market, probably because it turned out it's nice for things like this, but it's hardly in the 99% can't afford domain, especially if you're using it to replace a never-ending rent at which point it will pay for itself very rapidly, especially for heavy LLM users.

      In any case, we've gone from requiring supercomputers, to requiring very high end computers, to requiring $1600 video cards. It's tracking the exact same path that image rendering systems took (which if you haven't been keeping up there, now run excellently on pretty much any plain old computer), and we'll probably be there within a couple of years if not much sooner.

    • MrDrMcCoy 24 minutes ago
      Ternary Bonsai 2 might be for you.
  • quietFalcon 1 hour ago
    Nice, though generation speed is the easy half for MoE offload, what's your prompt processing look like at say 16k context?
  • 0xbadcafebee 26 minutes ago
    Lol, sure, if you quant it to hell (Q2) it'll go real fast...

    They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.

    It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model.