From the creator of Redis; run LLM locally with ds4

(dwarfstar.sh)

62 points | by fibo 2 hours ago

5 comments

  • twoodfin 17 minutes ago
    https://github.com/antirez/ds4

    The project GitHub page is a much better introduction for the hn crowd.

  • vlowther 24 minutes ago
    It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next (https://github.com/antirez/ds4/pull/1115 if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.
  • simoiacos 42 minutes ago
    Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.

    I'm also looking into expanding the protocol and the engine to support various steering techniques.

    https://github.com/simoneiacomino/xenolith

  • doctorpangloss 1 hour ago
    the problem is the dsv4 checkpoint so quantized isn't very good
  • 123-11292 55 minutes ago
    Cannot read the vibe coded website because it hangs and lags.

    So recently opposition to the local LLM narrative pops up here and there and we get reassurance immediately. Is this automated by AI now? Make a sentiment analysis and produce slop articles that suppress last week's opposition?