Markovian Thinking: AI’s New Frontier in Million-Token Reasoning

Oct 23, 2025 | AI

Breaking the Chains of Computational Doom

In the shadowy corridors of AI labs, a new technique named Markovian Thinking is causing a stir. Developed by Mila researchers, this method promises to transform large language models (LLMs) from lumbering giants into nimble thinkers. By cunningly breaking down reasoning into manageable chunks, these models can now engage in complex reasoning without the usual computational heart attack. Enter Delethink, the environment where this magic happens, slicing the reasoning chain into fixed-size bits and laughing in the face of computational doom.

The numbers are tantalizing. For a 1.5 billion parameter model, Markovian Thinking can slash training costs by over two-thirds compared to the usual methods. This isn’t just a cost-saving measure—it’s a paradigm shift. Instead of cutting corners or prematurely wrapping up processes, Mila’s approach embraces the full length of reasoning. It’s like giving your AI the stamina of a marathon runner without the carb-loading.

The Curse of Quadratic Growth—Exorcised

Long-chain reasoning has been the bane of AI developers, a quadratic curse that inflates computational costs faster than a balloon in a vacuum. Traditional methods, like LongCoT, have tried to stretch the AI’s reasoning muscles using reinforcement learning (RL). But there’s a catch: as the AI’s ‘state’ expands with each token, the costs skyrocket, rendering complex problem-solving a financial nightmare.

Mila’s team, however, decided to sidestep this quadratic quagmire. By designing an RL environment that keeps the context window constant, they’ve turned quadratic growth into a linear stroll in the park. The Markovian Thinker doesn’t just think longer; it thinks smarter, focusing on what’s essential. It’s a bit like teaching your AI to pack light for a long journey, ensuring it carries only the essentials from one reasoning chunk to the next.

This is not just theoretical. Delethink forces models to reason within 8,000-token chunks, resetting the context with each new chunk. This clever sleight of hand allows the model to retain crucial information without being bogged down by its own verbosity. It’s like giving your AI a memory that refreshes without losing the plot.

Delethink: The AI’s New Playground

In the grand experiment known as Delethink, researchers put their theories to the test with R1-Distill-1.5B, training it on competition-level math problems. The results? A model that could reason up to 24,000 tokens, matching or even surpassing its LongCoT-trained peers. On tasks like coding and PhD-level questions, Delethink didn’t just hold its ground; it danced circles around the competition.

The real kicker comes when scaling beyond the training budget. While LongCoT models hit a performance ceiling, Delethink-trained models kept improving, solving problems after reasoning through a staggering 140,000 tokens. This isn’t just a win for AI researchers; it’s a boon for enterprises looking to cut costs without sacrificing performance. Imagine training a model with a 96,000-token average thinking length in just 7 H100-GPU-months instead of 27. That’s not just efficiency; it’s a quantum leap.

The Future of AI: Thinking in Hyperdrive

The implications of Markovian Thinking extend beyond mere efficiency. This technique hints at a future where AI models can engage in reasoning for millions of tokens, opening doors to scientific discoveries previously thought impossible. It’s like giving your AI the ability to ponder the mysteries of the universe without needing a cosmic calculator.

Interestingly, off-the-shelf models already show a knack for Markovian reasoning, suggesting that this approach is not just compatible but complementary to state-of-the-art models. The researchers’ experiments with larger models like GPT-OSS 120B demonstrated robust performance, reinforcing the notion that Delethink is not just a tool but a catalyst for next-gen AI capabilities.

As the researchers boldly claim, Markovian Thinking could be the key to unlocking AI’s true potential, enabling models to ‘think’ for extended horizons. It’s a tantalizing glimpse into a future where AI doesn’t just solve problems—it redefines them. And in this brave new world of AI, who knows what discoveries await?

Scientific Facts Worth Knowing

  • •💡 Markovian Thinking can reduce training costs by over two-thirds for a 1.5B parameter model.
  • •💡 Traditional LongCoT methods face quadratic growth in computational costs.
  • •💡 Delethink allows models to reason in fixed-size chunks, improving efficiency.
  • •💡 Delethink-trained models can solve problems beyond their initial training token budget.
  • •💡 Markovian Thinking enables models to engage in extended reasoning for scientific discovery.