Nine coding harnesses vs. your laptop

(nasutton.notion.site)

96 points | by nasutton12 10 hours ago

13 comments

  • swiftcoder 12 minutes ago
    I'd be interested to see how Reasonix stacks up here - they seem to have spent a lot of effort on tuning prefix cache reuse
  • OleksandrC 3 hours ago
    If you're looking for a coding agent that would fit nicely into resource-constrained environments (such as laptops, or tiny VPS servers, or tiny single-board computers, etc), and would also work great with local models - you might also like hax (https://usehax.dev/). 0.7 MB dynamically linked native C binary, few MBs of RAM usage when running, auto-discovers config from running local llama-server, and uses minimalist system prompt and tools for lean context usage.
    • mischief6 2 hours ago
      ive looked at your project before but forgot about it. I think mine is in a similar vein: https://github.com/mischief/clm

      it grew out of annoyance of dependencies on js runtimes, probably similar to you. mine additionally works on solaris and esp32.

      could be interesting to collaborate!

  • larodi 40 minutes ago
    I can see this pattern of many people using Qwen 3.8 27B for local inference both on Apple Silicon and x86. This implies the model must be very good, given all these peoples' opinion converges on it.
  • julesrms 1 hour ago
    HN seems to have had a stream of agent harness benchmarks floating past. And every time I wonder where the people who create these tests are looking when they're deciding which harnesses to test? Because right now nobody seems to bother testing mine! (https://juggler.studio)

    I know Juggler's very new, but there's so much churn going on in this area that it's hard to know where I should be pushing it. It's hard to guess whether juggler's strengths would played well with a particular test like this, or made it look bad, all feedback about the kind of parameters people are interested in is useful to know when I'm deciding what to optimise.

  • alex_john_m 3 hours ago
    What is this supposed to mean?

    "it spreads up to 50% between nights, so nothing between the lean arms is a finding."

    • entrope 3 hours ago
      The same thing as the last word of "That is the difference between 22 and 226 seconds, measured." Techies I know would mostly omit "measured"; the rest would show, not tell.
      • arjie 1 hour ago
        Pangram fires as usual. Human opening, 75% machine.
    • nxobject 1 hour ago
      Whatever it is, it’s just as hilarious as Engrish…
    • stavros 2 hours ago
      Means Claude can't write for shit.
  • toasty228 2 hours ago
    A bit off topic because I'm not using local models, but I recently benchmarked codex vs pi vs omp with my workload and found codex to be both faster and more token efficient than pi/omp. There was not a single case for which pi was faster/cheaper
    • weiran 1 hour ago
      Pi is a very basic harness by design. On the other hand OMP is a bloated mess of other people’s workflows.

      The trick with pi is to extend it yourself as you use it. It’s pretty easy to do.

      • toasty228 1 hour ago
        Every single article and banchmark say pi saves token by default, the more I add extensions the more token hungry it gets.

        pi used 2-3x the tokens of codex. pi with subagent pkg used 8x-10x the tokens of codex.

        I don't see how adding bloat to pi would make it more token efficient if the baseline is so poor to start with

  • asdfsa32 36 minutes ago
    What is with the website though? Rubbish scrolling. Junky rendering with artifacts if you scroll fast.
    • embedding-shape 28 minutes ago
      Notion is a "knowledge database" with awful performance and jank, that some people have decided sounded like a perfect place to host their blog for whatever reason. But these shared pages been as buggy as the first time I saw them years ago, not sure what they're doing.
  • humbleferret 1 hour ago
    Nice writeup! I imagine these results change as harnesses are updated, so you'd need to frequently rereview.

    I'd love to see a tiny, reproducible benchmark repo that anyone can drop on their own hardware and then run against all harnesses at once to compare the per turn prefix token count, time to the first token, experienced tokens/sec (and prefill), cache reuse % and a pass rate on a deterministic set of small tasks. I think it could also be useful to have some way to share results and hardware for others to compare.

  • tontinton 2 hours ago
    I've made https://maki.sh for use cases such as this
    • nopurpose 10 minutes ago
      Do I understand correctly that in your harness models don't call tools and pass output of one to another via context, but instead code whole pipeline as small on-demand tools and see only final output?
  • teekert 3 hours ago
    Fun reference I tested on 32 GB ram laptop with no extra GPU: llama.cpp: “what is ls”, almost immediate starts answering at one ~word/sec. Ask opencode with same model (some gwen e4b or something) to check what’s in its working directory: 20 min to response.
  • montyanne 4 hours ago
    Neat article.

    “Chad” initially looked interesting but the minute I saw the ai-written markdown and giant commit I just left. I just can’t bring myself to read someone elses’ slop, regardless of performance.

    If all a developer hand writes is a truthy and readable markdown document, I really don’t care if the rest of the project is vibe coded, but I struggle to get interested in AI generated summaries and docs.

    • CGamesPlay 3 hours ago
      I for one am excited to learn more about how it spreads up to 50% between nights, and how nothing between the lean arms is a finding.
      • ramon156 1 hour ago
        Don't forget it's, measured.
    • montyanne 4 hours ago
      [dead]
  • NooneAtAll3 2 hours ago
    wtf is wrong with scrolling on that website?
  • wip0 4 hours ago
    [dead]