Skip to content
Stoa Labs

The Stoa Labs thesis

Last reviewed · July 2026

Explore, then engineer what holds

Agentic AI is young enough that most of its foundations are still unexamined. What lasts will belong to whoever examines them: the teams that explore from first principles, measure what is actually true, and engineer what holds.

Agentic AI is moving faster than its foundations can settle.

Every layer of the stack is still an open question. How an agent is scaffolded can change the outcome as much as which model runs inside it. How agents acquire skills and tools, how long-running work survives a crash or a deploy, what an agent should remember and what it should be made to forget, how anyone can tell whether a run actually worked, and how software itself should be written now that agents write so much of it: each of these is argued about loudly and measured rarely. The half-life of a confident assumption in this field is a few months.

The common response is to borrow. Teams inherit whatever framing arrived with their tools, and most opinions about what agents can and cannot do are borrowed assumptions wearing the costume of fact. Building on an unexamined foundation works right up until the work matters.

Our response is the company itself:

Stoa Labs exists to explore agentic AI from first principles, and to engineer what it learns into systems that can be trusted with real work.

We mean that sentence as a bet, not a decoration. It has two halves. Exploration comes first because in a field this young, understanding is the scarcest input: ideas get followed because they are worth understanding, before anyone asks what they are worth commercially. Engineering comes second because understanding that never becomes a working system is commentary. The two run as one loop: explorations produce findings, findings that survive scrutiny move into development, and what gets built sends new questions back into the research.

The exploration has a map. Stoa Labs maintains seven research areas across the agentic stack: harness engineering; agent extensibility and tooling; orchestration and durable execution; context, memory, and knowledge; evals, observability, and AgentOps; agentic-driven development; and frontier research. They are not seven products or seven promises. They are the places where the field’s foundations most need checking.

Across all seven, the hardest problems keep reducing to the same shape: we can generate and act, but we cannot cheaply verify. Harnesses need verifiers and stop rules. Skills need evaluation before they can be trusted. Long-running work needs proof that it did what it claims. Code written by agents needs review that does not consume the time it saved. When every direction of research runs into the same wall, that wall is the frontier. Verification is where our work concentrates.

Ideas graduate on evidence. An exploration is allowed to end with “this did not pan out”; that is a finding, not a failure. What survives becomes development: methods, systems, and eventually products, each claiming only the evidence state it has actually reached. We will publish what we measure, including what did not work, because the strongest claim we can make is one that can be inspected.

One conviction holds across the map: autonomy is earned per capability, not granted per agent. Authority belongs where the evidence supports it, and like everything else in the lab, that discipline keeps its place only as the evidence allows, with our own systems as its first proving ground.

We run the lab the way we describe it. Stoa Labs is a deliberately small team of engineers building its own operations on the systems it studies, because a small team of good engineers using agents can do great things, and because running on what you research is the cheapest source of honest evidence there is. A map this wide is honest work for a small team only because agents carry so much of it, and that arrangement is itself part of what the map is testing.

The Stoic connection is not that ancient philosophy anticipated software agents. It is a shared discipline of attention: distinguish what can be governed from what cannot, then act deliberately within that distinction. The Stoa Poikile was a covered walkway in Athens where Zeno taught in public, in the middle of working life. Stoa Labs takes the stoa as a model for technical inquiry conducted openly and close to working life, with claims revised when the evidence changes.

So that is the bet. The teams that create lasting value with agentic AI will not be the ones with the most confident assumptions or the most autonomous demos. They will be the ones who checked the foundations, measured what they built, and engineered what holds. That is the lab we are building.

The research map: seven areas How the lab works