Subconscious raises $5.1M to build the inference platform for long-running agents | exclusive

Become a member of GB MAX to gain exclusive access to the industry and to the most influential global B2B leadership community in the business of gaming, entertainment, and tech. Join now and also get a VIP ticket to GamesBeat Next (Nov 2-3, SF).

Subconscious said it has raised $5.1 million in funds to launch an inference platform designed for long-running agents.

It’s build on this notion: MIT researchers discovered a way to dynamically compress 95% of an AI agent’s context and make it much more efficient and accurate.

Now, they’re turning that core technology into an opinionated inference platform to power agents that run faster, for longer, at a lower cost with no changes to the underlying hardware, AI models, or apps that use them.

MassVentures led the round, with participation from Foothill Ventures, Underscore VC, E14 Fund, and the Agent Fund among others. The inference platform is available today for developers, with an on-prem deployment package available for enterprises.

Language models are used in many ways: as a chatbot, as a classifier, as a document writer, or as an AI agent. Among all these uses, agents are extremely computationally intensive and require processing thousands of times as many tokens over long periods of time. Agents are the most expensive way to use AI models, but they generate the most valuable work. Already they’ve transformed software engineering, and they’re growing in usage among salespeople, lawyers, scientists, marketers, consultants, and all kinds of knowledge work. Subconscious believes agents will make up virtually all inference in the very near future.

Subconscious built an inference platform to take advantage of the unique challenges in powering long-running agents. Born out of MIT research into inference, the system uses dynamic context compression and highly efficient caching for a stepwise gain in performance. For agents that consume beyond 200k tokens, Subconscious can generate tokens faster, extend the effective context window of models to 5m+ tokens, improve performance on key agentic benchmarks like coding and workflows, and decrease costs by up to 80%.

On the TriE benchmark, a benchmark built to measure system performance, Subconscious’s inference runtime generated tokens 3.5 times faster than SGLang on coding tasks and supported 2.3x as many concurrent requests. Handling more concurrent requests means squeezing more work out of GPUs, effectively doubling or tripling the size of a cluster.

On DeepSWE, a benchmark built around long coding tasks, GLM 5.2 hosted on Subconscious solved 46 percent of problems at an average cost of $2.79. The same model served on standard inference infrastructure scored 44 percent and cost $3.92. Compression and caching from Subconscious improve both the model’s efficiency and its accuracy.

Subconscious is live in production today powering coding agents and agentic products. In July, a 20-person engineering team switched from using Claude to using the GLM 5.2 model hosted on Subconscious to power their coding agents. In the two months since, they’ve cut their monthly AI spend from $40,000 to $6,000, their engineers report faster token throughput, and they have yet to hit their rate limits.

One engineer on the team ran an extremely long agent trace across 4,571 turns and 9,556 tool calls, and Subconscious recorded 449 million tokens where a conventional runtime would have billed 2.6 billion. Just as important is what the engineering team didn’t report: despite aggressive context compression, they reported no loss in model capability.

As AI costs have skyrocketed many companies have moved to cut costs, but engineers want more agents running for longer periods of time on harder problems. Subconscious allows companies to have it all.

“Open models finally got good enough this summer that their quality vs closed source models stopped being a compromise,” said Jack O’Brien, CEO of Subconscious, in a statement. “Meanwhile every engineering team I talk to has put a ceiling on what it spends per developer. Teams need to spend less, but engineers are addicted and there’s no going back. We run open models in a way that’s enhanced for coding agents, so teams can have it all.”

I asked O’Brien what the inspiration was for starting the company. He replied, “My co-founder Hongyin Luo and I had been building roughly the same thing in parallel for two years. We kept coming back to agents: they were inefficient and hard to build reliably, even with the hype through the roof. So we rethought what an inference stack would look like if it were built from the ground up to serve agents.”

As far as the competition goes, O’Brien said, “Most inference companies build one-size-fits-all systems for chats, one-shot requests, and agents alike. But agents are much harder to serve well. We built our whole stack around long-running agents, and our core technology comes from novel MIT research, which gives us a year-plus head start.”

The Subconscious inference platform is live today, and teams can start powering coding agents like Claude Code, Codex, Pi, and OpenCode with Subconscious in about 30 seconds. Teams can also deploy the inference system on their own GPUs for maximum cost savings, fully air-gapped data protection, and ultimate control.

“The Subconscious team is the best team on the planet to solve one of AI’s toughest challenges,” said Stacy Swider, VP of Investments at MassVentures. “We could not be more excited to be a part of their growth as they scale up to support thousands of companies and billions or even trillions of AI agents.”

Subconscious is an inference platform designed for long-horizon agents. Engineers choose Subconscious to power their agents to finish complex tasks faster, run for longer, and lower their costs substantially. The company is based in Cambridge, Massachusetts, and it has more than 10 people.