Studio notes · jiongan mu

Subagents
as code.

Most harnesses use subagent profiles or let the main agent construct a subagent by dynamically assigning its system prompt and tools. Graph engineering adds explicit coordination. But treating the subagent itself as executable code is both simpler and much more powerful.

Aug 01, 2026 · 4 min read · ~830 words

Most agent harnesses implement one of two subagent patterns right now: pick a subagent profile, or construct a subagent at runtime.

The first is what you see in tools like Claude Code. The harness has a catalog of subagent profiles with fixed prompts and tool access, and the parent decides which one to call. The second lets the main agent construct a subagent on the fly by dynamically assigning a system prompt, a scoped subset of tools, and sometimes a model for the task. I wrote about both patterns in Subagent designs.

The shift to graphs

“Graph engineering is designing agentic systems as explicit graphs instead of implicit loops. Nodes are units of capability. Edges are decisions. State is an object with a schema, checkpointed every time you cross an edge.” — Josh C. Simmons

This is essentially the idea behind Ensemble, which I started working on in April: orchestrating agent teams as an explicit graph rather than hiding coordination inside one agent loop. Give it a try—you can shape a custom graph around the job, whether that is coding, finance, research, or something else entirely.

Make the subagent code

I think there is a simpler primitive hiding underneath all of this: let the main agent supply Python code that can spawn agentic loops.

Tool contract

The main agent calls invoke(name, description, code). Those three arguments define the tool:

tool definitionJSON
{
  "name": "research_brief",
  "description": "Researches and summarizes one topic",
  "code": "def run(inputs, ctx):\n    ...\n    return {'brief': result}"
}

The code field looks something like this:

generated nodePython
def run(inputs, ctx):
    result = ctx.agent(
        prompt="Research current subagent orchestration patterns and return a concise, sourced brief",
        tools=["web_search", "web_fetch"],
        label="researcher",
    )
    return {"brief": result["content"]}

Instead of adding another special subagent schema, the harness receives a Python function. That function can call ctx.agent(prompt=..., tools=..., label=..., model=...), then transform or use the result however it wants. The optional model argument can route the spawned loop to a specific LLM. Spawning subagents becomes an operation available inside code, not a separate orchestration primitive.

Once subagents are ordinary function calls, the rest is just Python. Use for or while loops to spawn them iteratively, use concurrency libraries to invoke multiple agentic loops in parallel, define the returned result programmatically, and pre-process inputs or post-process LLM responses before passing anything onward.

A concrete example

Say the request is:

“Build a Markdown report covering every company in Y Combinator’s Spring 2026 batch. For each company, include its name, a concise description, its founders, and a short summary of each founder’s background.”

Without subagents

One agent has to discover the batch, research all 196 companies, track every founder, and assemble the report inside the same context.

Single-agent loopone context · 196 companies
First stepfind 196 company names
↓
repeat × 196 · same context
search company
↓
read sources
↓
summarize company + founders
↓
next company ↻
↓
inside the repeated loop · context window fullcompact earlier history

This is going to be slow, and repeated compaction across the 196-company loop can degrade the quality of the final output.

Typical subagent implementations

Most implementations give the main agent a company-specific subagent tool call:

one company-specific invocationtool call
invoke_subagent(
    task=(
        "Research [company name]: explain what the company does, "
        "identify its founders, and summarize their backgrounds."
    )
)

The subagent prompt can stay the same; only the company name changes. Whether the subagent uses a fixed profile or is configured dynamically, the main agent still has to invoke the tool 196 times—once per company—then track every returned summary and assemble the report. The question is whether it can reliably issue all 196 invocations in the first place, without skipping companies or stopping early.

Tool-mediated delegationmain agent + invoke_subagent
main agentprepare 196 company names + one shared research prompt
↓
196 tool calls
return 196 summaries to the main agent
↓
main agentvalidate + assemble the report

Even if the harness can run the calls in parallel, the main agent still has to generate 196 separate tool invocations, rewriting the task and other invoke_subagent arguments for every company. Producing that many calls can take a long time before the subagents even begin their work.

Our approach: subagents as code

The main agent can instead generate one Python function that discovers the batch, fans out one research loop per company, and assembles the returned sections into a single Markdown file:

parallel company researchPython
import json
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path


def run(inputs, ctx):
    # Discover the batch once before spawning company research.
    batch = ctx.agent(
        prompt=(
            "Find all 196 companies in Y Combinator's Spring 2026 batch. "
            "Use the official YC directory as the primary source. "
            "Return only JSON in this exact shape: "
            "[{\"name\": \"Company\", \"profile_url\": \"https://...\"}]."
        ),
        tools=["web_search", "web_fetch"],
        label="yc-s26-index",
    )
    companies = json.loads(batch["content"])

    def research(company):
        # Give each company its own isolated agent loop.
        result = ctx.agent(
            prompt=(
                f"Research {company['name']} using its YC profile and other "
                "reliable sources. Return only Markdown in this exact format:\n\n"
                "## {company name}\n\n"
                "{one-sentence description}\n\n"
                "**Founders**\n"
                "- [Founder name](source URL): short background\n\n"
                "**Sources**\n"
                "- [Source title](source URL)"
            ),
            tools=["web_search", "web_fetch"],
            label="yc-company-researcher",
        )
        return result["content"]

    # Run up to eight company research loops at the same time.
    with ThreadPoolExecutor(max_workers=8) as pool:
        sections = list(pool.map(research, companies))

    # Assemble the report, write it to a file, and return its path.
    report = "# YC Spring 2026 companies\n\n" + "\n\n".join(sections)
    output_path = Path("yc-spring-2026.md")
    output_path.write_text(report, encoding="utf-8")
    return str(output_path)

This is clean. You can add deterministic checks or an LLM judge after any step. Because the orchestration is code, you can validate, retry, route, transform, or persist the work however you want.

Running it for real

The main adoption concern is runtime. A program that launches many agentic loops can run for minutes—or longer—so executing it inline would block the agent from responding to new user messages.

The tool should start the program as a background job and return immediately. When the job finishes, it should emit a completion event into the conversation that triggers another LLM turn.

Long-running agent program non-blocking lifecycle
01 · STARTLaunch sandboxed Pythonbackground job
02 · RETURNGive control back immediatelyconversation stays live
03 · EVENTEmit completion signalresult attached
04 · RESUMETrigger another LLM turncontinue with output

Running this safely also requires an actual sandbox for the generated Python. Cost needs guardrails too, although many of the spawned agentic loops can use cheaper models. One extension I am considering is letting the Python program define tools as ordinary functions. Those functions could then be passed into ctx.agent(...), where the spawned agent can call them just like any other tool.

Try it out in Ensemble.