Exploration
Ideas get followed because they are worth understanding.
Research is the center of this company. The lab maintains seven research areas across the agentic stack: how agents are built, extended, coordinated, given state, measured, secured, and used to write software, and the open questions beyond. The exploration is open-ended on purpose, the engineering is where ideas prove out, and what the lab learns flows into the consulting practice, never the other way around.
Areas 01 through 07
The map, explored in the open.
Seven areas cover the stack an agentic system actually lives on. Each is a claim of active research attention with its evidence stated honestly, not a product or a promise.
Harness Engineering
An agent is a model plus a harness, and the harness gap is measured, not anecdotal. We study how scaffolding changes what an agent can do and how far it can be trusted.
Agent Extensibility & Tooling
Agents are only as capable as the skills, tools, and delegation structures around them, and published capabilities often weaken under measurement. We study how agents acquire capabilities and how those capabilities are evaluated, secured, and governed.
Orchestration & Durable Execution
A workflow that cannot recover from a crash, deploy, or interrupted approval is not ready for consequential authority. We study how long-running and multi-agent work stays correct.
Context, Memory & Knowledge
Agent behavior depends on the state it receives: context can bloat, memories can go stale, and retrieval can return the plausible instead of the true. We study agent state as an engineering discipline.
Evals, Observability & AgentOps
The hardest question in agentic AI is not “can we build it” but “did it work.” We study outcome-level evaluation, judge reliability, and the operational signals that say when trust should contract.
Agentic-Driven Development
AI can write code; the open question is what evidence justifies trusting the result. We study verification-first development and the authority coding agents should receive in consequential repositories.
Frontier Research
Every agentic system embeds a bet about where its logic should live: in prompts the model interprets, or in code the runtime enforces. The lab’s open-exploration area.
Explorations here become public claims only with evidence, and every published claim carries its method, context, and limits. Several areas anchor our service families with independently developed methods; client work never enters the research.