11X Cheaper ChatGPT: A Pathway (AI lab focused on Post-Transformer ideas) released BDH-CQ, a 150M-parameter reasoning model, and its benchmark results claim 11x lower ope than OpenAI’s ChatGPT-style reasoning paths—without forcing the AI to “think out loud.” The inciting moment for this shift wasn’t a flashy demo, but Pathway’s publication of ARC-AGI-1 benchmark outcomes that tried to measure Intelligence per dollar, not just raw capability.
2026: Pathway publishes BDH-CQ results that target “token cost” reasoning
In the early 2020s, “chain-of-thought” style intermediate reasoning became a popular recipe for stronger large language model behavior. But Pathway’s BDH-CQ sets a different goal: keep reasoning internal instead of gene
Worth noting: the company’s argument is not that reasoning is unnecessary—it’s that the way you externalize reasoning can charge users for tokens you don’t actually need. Pathway’s CEO and co-founder Zuzanna Stamirowska told TechRadar that traditional AI “pays a steep token cost for reasoning” due to design choices, not any law of Intelligence.

Key Details: BDH-CQ hits 29.5% pass@2 at $0.0007 per task
The strongest concrete datapoint in the announcement is performance on the public ARC-AGI-1 evaluation set. TechRadar reports BDH-CQ reached 29.5% pass@2 with a model size of 150 million parameters. That’s paired with an asserted computed inference cost of $0.0007 per task.
Here’s the thing: the comparison is framed as about ope11 times cheaper than ChatGPT’s underlying model, using BDH-CQ’s claimed inference economics as the main benchmark lens.
That said, the article emphasizes what’s happening under the hood. The BDH-CQ approach, as described in TechRadar’s write-up, performs reasoning internally rather than outputting lengthy intermediate text—effectively reducing the “talking during thought” overhead that can inflate token bills.
Context: why “thinking out loud” became standard—and why this challenges it
Big reasoning gains in LLMs often came alongside visible intermediate steps, which helped improve problem-solving and made the model’s process easier to inspect. In practice, though, these outputs can inflate the amount of generated text required per question. Pathway’s message is that “thinking out loud” isn’t inherently Intelligence; it can be a convenient behavior pattern that carries a hidden cost.
This is where we connect the dots to broader AI conversations. OpenAI’s technical publishing has repeatedly highlighted the value of reasoning and training approaches, including how models behave under different prompting and training regimes (see OpenAI’s research updates at the OpenAI Blog). But Pathway is pushing a counterpoint: you can restructure computation so the model reasons without producing as many intermediate tokens.
Worth noting: TechRadar also quotes Amazon Web Services as believing the BDH-CQ results are a promising step toward deploying advanced AI reasoning in real products. That matters because cost-to-serve and latency are often what gate “reasoning AI” from prototypes to everyday tools, even when benchmark scores look promising.
What’s Next: cheaper reasoning could reshape agent economics in 2026
If Pathway’s cost and accuracy claims hold up across more evaluations, the implication is straightforward: smaller models with internal reasoning pipelines could deliver competitive performance per query. That changes procurement conversations for AI features that rely on multi-step problem solving—customer-facing chat isn’t the only place reasoning matters.
In practical terms, teams building AI assistants can expect renewed focus on “Intelligence per token,” not just “Intelligence per parameter.” If internal reasoning reduces generated text, then inference-time compute and billing pressure can shift—freeing budgets for bigger context windows, retrieval, or tool use, rather than paying the model to narrate its steps.
That forward path also affects how developers choose runtimes and architectures. Instead of assuming that stronger reasoning requires more verbose intermediate outputs, the 2026 takeaway is to treat architecture as an economic lever—then benchmark not only correctness, but per-task compute and cost.
Related Articles
FAQs
What is BDH-CQ, and why is it being compared to ChatGPT?
BDH-CQ is a 150M-parameter reasoning model from Pathway, evaluated on ARC-AGI-1. TechRadar’s coverage compares its reported $0.0007 per task inference cost and 29.5% pass@2 results to ChatGPT’s reasoning economics, claiming about 11x cheaper operation.
Does the model still reason if it doesn’t “think out loud”?
Yes—Pathway’s approach, as described in TechRadar, keeps reasoning internal rather than gene
Why does token cost matter for AI products?
Because many systems bill users (or companies) based on generated tokens and inference workload. If reasoning can be computed with fewer generated intermediate tokens, then cost-to-serve can drop even when the task remains multi-step.
Was this article helpful?
Your feedback directly improves future articles on this site.





