Stealing Reasoning Traces from Proprietary LLM APIs

(stolen-thoughts.com)

46 points | by quantumgarbage 1 hour ago

9 comments

  • x312 1 minute ago
    Super cool that this works. I'm surprised these companies re-use the same encryption key across models!

    I wonder if you can use these for attacks, like this previous paper showing that if you know how a model reasons, you can "fake its thinking" to control it? https://news.ycombinator.com/item?id=48631888

  • Groxx 17 minutes ago
    >We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, ...

    Ha! I've been wondering if replaying across models would work, ever since https://blog.cryptographyengineering.com/2026/05/29/fooling-...

    I'm honestly rather curious if this was intentionally allowed, it's the sort of validation that's easy to miss (particularly if you're wading into the vibe waters). Seems like something that'd be absolutely riddled with possibilities for shenanigans.

    • yojo 5 minutes ago
      If you didn’t allow it, you wouldn’t be able to change models in the same conversation, as key parts of the context would be lost.

      Wouldn’t surprise me if the providers just remove that ability and lock the model once the conversation starts.

  • nervai 6 minutes ago
    Really cool work, you get the actual traces. Looks like the vendors can all reliably fix this one though.

    A harder to defend against approach here where they work backwards from the results and ask the model to generate a plausible trace: How to Steal Reasoning Without Reasoning Traces https://arxiv.org/pdf/2603.07267

  • iamcoder18 4 minutes ago
    This proves that OpenAI models reason in grug speak to save tokens! I wonder if open models are going to start doing that too to save on reasoning tokens.
  • fractorial 17 minutes ago
    Fascinating approach; however, a nightmare to scroll on mobile.
  • dboreham 6 minutes ago
    Can someone tell us how they were able to decrypt the encrypted payload? The article says they inserted the cyphertext into a session with a different model. Ok, but how does that allow you to decrypt it?
  • alansaber 10 minutes ago
    Neat.
  • locitra 2 minutes ago
    [flagged]
  • quantumgarbage 1 hour ago
    Proprietary reasoning can be recovered from its encrypted traces. Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext, without ever attacking the stronger model directly or triggering its anti-distillation safeguards.
    • the_af 8 minutes ago
      Why do you restate the abstract? Anyone can read it from the link.
      • ronsor 5 minutes ago
        This is Hacker News. You know people don't follow links and read.
      • mschuster91 3 minutes ago
        People don't read no links no more