Skip to content

AutoAgent Guide

AutoAgent is JarvisCore's autonomous reasoning profile. A minimal agent needs three class attributes: role, capabilities, and system_prompt. Production agents can also describe each capability, declare its authorized effects and systems, validate output, select an execution role, and request Nexus-backed credentials. The framework then handles LLM selection, Kernel routing, sandboxed execution, repair, and peer participation. For multi-step goals, goal_oriented = True enables Plan, Execute, Evaluate with automatic replanning.


Defining an AutoAgent

agents/researcher.py
from jarviscore import AutoAgent

class ResearcherAgent(AutoAgent):
    role = "researcher"
    capabilities = ["research", "synthesis", "web-search"]
    system_prompt = """
    You are a rigorous research analyst. Prioritise primary sources,
    cross-reference findings, and structure outputs as actionable intelligence.
    Always store your final output in a variable named `result`.
    """

The framework requires role, capabilities, and system_prompt; construction raises ValueError when one is absent. That is the smallest valid AutoAgent, not the complete production contract. Add the optional declarations below when the Mesh or Kernel needs them to route, authorize, or validate work.

Class Attributes

Attribute Required Description
role Yes Slug used for peer discovery, profile loading, and workflow routing
capabilities Yes Tags for capability-based peer discovery
system_prompt Yes Base LLM system prompt; framework raises ValueError if absent
description No One-sentence purpose used by peers for routing decisions
capability_descriptions No Mapping from capability name to a concrete routing description. Distributed planning uses it instead of guessing from a short tag.
capability_contracts No Mapping from capability name to authorized effects, provider systems, and optional planner-visible produces artifact descriptions.
output_schema No Pydantic model class enforced on CoderSubAgent execution output. Validate other role outputs or richer work products in an application subclass.
default_kernel_role No Preferred fallback role for specialist agents. Built-ins are "researcher", "coder", "communicator", and "browser"; products may register custom roles through an extended Kernel. Leave unset for generalists.
goal_oriented No Defaults to False; set True for multi-step goal decomposition
requires_auth No Defaults to False; set True to receive Nexus-backed _auth_manager

Production Mesh declaration

The class remains an ordinary AutoAgent subclass. The additional attributes make its routing and authority explicit:

agents/repository_reviewer.py
from jarviscore import AutoAgent


class RepositoryReviewer(AutoAgent):
    role = "repository_reviewer"
    description = "Reviews repository changes against correctness evidence."
    capabilities = ["code_review"]
    capability_descriptions = {
        "code_review": "Inspect a change, identify defects, and propose review findings.",
    }
    capability_contracts = {
        "code_review": {
            "effects": ["read", "propose"],
            "systems": ["github"],
            "produces": "ReviewReport(findings, inspected_revision)",
        },
    }
    default_kernel_role = "coder"
    requires_auth = True
    system_prompt = """
    You are a rigorous code reviewer. Ground every finding in repository
    evidence and return only review findings supported by the inspected change.
    Always store the final output in `result`.
    """

effects may contain read, propose, write, notify, or destructive. systems names the provider boundaries the capability may use. These contracts do not grant credentials by themselves: requires_auth controls credential injection, and provider policy still governs each operation.


Running an AutoAgent

Standalone script

main.py
import asyncio
from jarviscore import Mesh
from agents import ResearcherAgent

async def main():
    mesh = Mesh()
    mesh.add(ResearcherAgent)
    await mesh.start()

    results = await mesh.workflow("research-001", [
        {"agent": "researcher", "task": "Summarise the state of AI hardware in 2026"}
    ])

    step = results[0]
    print(step["status"])    # "success"
    print(step["payload"])   # the value stored in `result` by the generated code
    await mesh.stop()

asyncio.run(main())

Mesh() takes no mode argument. It auto-detects available infrastructure at start() time. Pass REDIS_URL in your environment and the Mesh will use Redis automatically. Pass it explicitly via the config dict if you need to override:

mesh = Mesh(config={"redis_url": "redis://localhost:6379/0"})

For P2P between nodes, add p2p_enabled: True and bind_port to the config dict. No mode string required.

FastAPI service

main.py
from contextlib import asynccontextmanager
from fastapi import FastAPI
from jarviscore import Mesh
from agents import ResearcherAgent

mesh = Mesh()

@asynccontextmanager
async def lifespan(app: FastAPI):
    mesh.add(ResearcherAgent)
    await mesh.start()
    yield
    await mesh.stop()

app = FastAPI(lifespan=lifespan)

@app.post("/research")
async def research(request: dict):
    results = await mesh.workflow("req-001", [
        {"agent": "researcher", "task": request["task"]}
    ])
    step = results[0]
    return {"status": step["status"], "output": step.get("payload")}

What workflow() Returns

mesh.workflow() returns a list of step result dicts in the original declared order. Each result carries the agent envelope plus its stable step_id:

Key Description
status "success", "failure", "yield", or "hitl" for a paused goal-oriented execution
output Programmatic task result
payload Alias of output on the standard Kernel path
result_summary Guaranteed plain-prose display summary
error None on success; failure or yield explanation otherwise
tokens Input, output, and total token counts
cost_usd Estimated execution cost
step_id Stable workflow step identity
goal_execution Present for goal-oriented planning or a classified direct Kernel turn

Always check step["status"] == "success" before consuming step["output"]. A failure has error plus a human-readable result_summary; a yield or HITL result carries continuation details in yield_metadata.

When calling agent.execute_task() directly, the returned envelope always carries result_summary: plain prose suitable for display, never JSON. Failure envelopes put a human-readable error sentence there. Structured data stays in output / payload / goal_execution for programmatic consumers.


Lifecycle Hooks

AutoAgent.setup() is called once by the Mesh after instantiation. It initialises the Kernel, LLM client, sandbox, FunctionRegistry, and loads the agent persona from YAML if configured. Override it for one-time setup:

async def setup(self):
    await super().setup()   # must come first
    self.data_client = await MyDataClient.connect()

teardown() is called during Mesh shutdown. Release any resources opened in setup():

async def teardown(self):
    await self.data_client.close()
    await super().teardown()

You will rarely need to override execute_task(). The Kernel pipeline routes internally. Override it only if you need to enrich the task dict before the Kernel processes it:

async def execute_task(self, task: dict) -> dict:
    enriched = {**task, "task": f"{task.get('task', '')}\n\nUser context: {self.user_id}"}
    return await super().execute_task(enriched)

System Prompt Best Practices

The system prompt is the primary way to shape the code the agent generates.

system_prompt = """
You are a financial data analyst.

Data source: Yahoo Finance API at https://query1.finance.yahoo.com/v8/finance/chart/{ticker}
Parse the JSON response carefully and handle missing keys with .get().

Your output must be stored in a variable named `result` as a dict with keys:
  ticker    (str)
  price     (float)
  change_pct (float)
  analysis  (str, 2-3 sentences)
"""

Tell the agent exactly what variable name to use (result), the expected shape of that variable, what APIs or tools are available in the sandbox, and what to do when data is missing or a request fails. Vague prompts produce fragile code.


Kernel Role Routing

The Kernel classifies each task and routes it to the appropriate sub-agent. You do not configure this directly. The Kernel infers the right sub-agent from the task. On specialist agents where every task always uses the same sub-agent, skip classification entirely by setting default_kernel_role:

class SlackNotifier(AutoAgent):
    role = "notifier"
    default_kernel_role = "communicator"   # always sends, never codes or searches

The four sub-agents the Kernel routes to are the CoderSubAgent for tasks that require writing and running Python, the ResearcherSubAgent for tasks that require web search and synthesis, the CommunicatorSubAgent for tasks that require formatting and delivering output, and the BrowserSubAgent for web navigation tasks when BROWSER_ENABLED=true is set.

Analysis, not code: the single_response contract

The answer must contain visible text and must not carry an explicit incomplete, filtered, refused, tool-call or unrecognized terminal reason. Otherwise the result uses status="failure", retaining any partial output/payload, finish_reason, provider_metadata, provider/model, tokens and cost for diagnosis. It never retries or reports a new warning status. Nonempty responses from older/custom clients with no finish reason remain compatible; token counts alone never establish truncation.

To set an output budget for this turn only, include a positive integer max_output_tokens in the execution contract, for example: {"execution_shape": "single_response", "max_output_tokens": 8192}. This is forwarded as generate(max_tokens=8192). Invalid values fail before an LLM call. Omitting it retains the existing configured provider budget; model limits still apply. A reasoning model may spend part of that budget on internal reasoning rather than visible text.

Many agent tasks need exactly one LLM completion: render the system prompt, ask the question, return the answer. No planner, no routing, no code generation. Declare this shape per task with an execution contract:

results = await mesh.workflow("analysis-001", [{
    "agent": "market_analyst",
    "task": "Analyse EURUSD H1 and return your thesis as JSON.",
    "context": {"execution_contract": {"execution_shape": "single_response"}},
}])

With single_response declared, execute_task() runs one completion against the agent's system prompt (including its persona profile, if one is loaded) and returns the standard result envelope with token and cost telemetry. The Kernel pipeline is skipped entirely, so an analysis prompt never reaches the Coder sub-agent and never produces TOOL/DONE protocol errors.

Use this for tasks where the whole job is the answer: analysis, classification, extraction, drafting. When the task needs tools, search, or several turns of reasoning, drop the contract and let the Kernel route it.


Coder Sandbox

When the Kernel routes a task to the CoderSubAgent, code executes inside CoderSandbox: a deliberately file-capable execution environment that is distinct from the SandboxExecutor used by other sub-agents.

SandboxExecutor (used by ResearcherSubAgent and CommunicatorSubAgent) blocks open(), subprocess, and filesystem access. CoderSandbox intentionally grants those capabilities, scoped to a controlled workspace directory.

Inside every execution, generated code receives these names in its namespace:

Name Type Description
workspace Path Project root directory. All file reads and writes are expected here.
output_dir Path workspace/output/: the preferred write location for produced files.
blob_path(name) Path Shorthand for output_dir / name. Creates parent directories.
bash(cmd) BashExecutor Run allowed shell commands. Returns {success, stdout, stderr, returncode}.
git GitHelper High-level git: checkout_branch, add_all, commit, push, describe_pr.
nexus_call async fn Make authenticated HTTP calls to any provider via Nexus. Never exposes credentials.
json, os, re, Path, datetime, math, shutil, uuid stdlib Pre-imported for convenience.

The Coder can import any installed package from within the sandbox. The security boundary is the bash allow-list, not Python builtins.

The bash allow-list covers: git, pip, pip3, cp, mv, mkdir, rm, ls, cat, echo, touch, find, grep, sed, awk, pandoc, npm, npx, node, python, python3, curl. Hard-blocked regardless: sudo, rm -rf /, eval, curl | bash, and pipe-to-shell patterns.

The CoderResult output shape: what gets written to result by generated code and surfaced in step["payload"]:

result = {
    "success":        bool,         # did the task complete?
    "files_created":  [str],        # absolute paths of files written
    "files_modified": [str],        # absolute paths of files changed
    "git_branch":     str | None,   # branch name if git ops were performed
    "stdout":         str,          # captured print() output
    "data":           Any,          # any structured data to return
    "error":          str | None,   # error message if success=False
}

Configure the workspace root and timeouts:

.env
# Workspace directory: defaults to the process working directory
CODER_WORKSPACE=/path/to/your/project

# Sandbox execution timeout in seconds (default: 300)
SANDBOX_TIMEOUT=300

When writing system prompts for agents that route to the Coder, always tell the agent what files it should write and where. The Coder will use output_dir by default if you do not specify a location.


Execution Budgets

Every sub-agent dispatch runs against an ExecutionLease: a token, turn, and wall-clock budget specific to the sub-agent role. When the lease is exhausted, the Kernel stops the sub-agent and returns whatever has been produced so far.

Role Thinking tokens Action tokens Total tokens Wall clock Turn fuse
coder 132,000 108,000 240,000 4 min 32 turns
researcher 180,000 60,000 240,000 4 min 36 turns
communicator 72,000 48,000 120,000 4 min 18 turns
browser 60,000 60,000 120,000 5 min 28 turns

The turn fuse is an emergency hard-stop. If a sub-agent reaches 32 turns without completing, the Kernel terminates it regardless of token budget. This prevents runaway loops on adversarial or ambiguous tasks.

You do not configure lease budgets per-agent: they are role-level defaults. If a specific task consistently exhausts the budget, the right fix is usually to narrow the task scope (break it into smaller steps in the workflow DAG), not to increase the budget.

Token consumption appears in step["tokens"] on every AutoAgent workflow result:

results = await mesh.workflow("report", [
    {"id": "analyse", "agent": "analyst", "task": "..."}
])
step = results[0]
print(step["tokens"])
# {"input": 14200, "output": 3100, "total": 17300}

Model Routing

JarvisCore routes each sub-agent dispatch to a capability tier (nano, standard, or heavy) based on the role and an optional per-step hint. Pass complexity in a workflow step to override the default for that dispatch:

results = await mesh.workflow("task-001", [
    {"agent": "analyst", "task": "Summarise this paragraph.",       "complexity": "nano"},
    {"agent": "analyst", "task": "Produce a competitive analysis.", "complexity": "heavy"},
])

When complexity is omitted, the role's built-in default applies: communicator defaults to nano, researcher and browser default to standard. See Model Routing for the full tier configuration, fallback chain, and provider compatibility reference.


Infrastructure Injection

The Mesh injects stores, mailbox, and HITL before setup() runs. Peer clients and optional authentication are attached later in mesh.start() and are available before the Mesh accepts work.

Attribute Available when
self._redis_store REDIS_URL is set
self._blob_storage Always: falls back to local filesystem
self.mailbox Always: in-memory by default; Redis-backed when configured
self.hitl When the framework HITL component is available; Redis-enhanced when configured
self.peers After mesh.start() attaches local or distributed peer transport
self._auth_manager After setup(), when requires_auth = True and authentication is configured
from jarviscore.memory import UnifiedMemory

async def setup(self):
    await super().setup()
    self.memory = UnifiedMemory(
        workflow_id="wf-001",
        step_id=self.role,
        agent_id=self.role,
        redis_store=self._redis_store,
        blob_storage=self._blob_storage,
    )

Agents running inside a workflow write their step output to step_output:{workflow_id}:{step_id} in Redis. Read a prior step's output from another agent using RedisMemoryAccessor:

from jarviscore.memory import RedisMemoryAccessor

accessor = RedisMemoryAccessor(self._redis_store, workflow_id="wf-001")
raw = accessor.get("fetch")
prior = raw.get("output", raw) if isinstance(raw, dict) else {}

Checkpointing

Checkpointing happens automatically. After every OODA loop turn, the Kernel calls UnifiedMemory.save_checkpoint() which writes the current KernelState (all accumulated findings, tool history, and reasoning progress) to Redis under checkpoint:{workflow_id}:{step_id}. If the process crashes and the workflow restarts, the WorkflowEngine detects the existing checkpoint and resumes from that turn rather than restarting from scratch.

You do not call save_checkpoint() or load_checkpoint() in application code. The framework manages this transparently when REDIS_URL is configured. Without Redis, there is no persistence and no crash recovery: the workflow starts from the beginning on failure.


Nexus Auth: requires_auth

Set requires_auth = True on agents that call connected third-party services. When connected-app authentication is configured, the Mesh creates an AuthenticationManager and injects it as self._auth_manager after setup(). The Kernel exposes nexus_call inside the sandbox; generated code never receives the manager or raw credentials.

class GitHubAgent(AutoAgent):
    role = "github_agent"
    capabilities = ["github", "code-review"]
    requires_auth = True
    system_prompt = """
    Use the available Nexus-backed GitHub tools for approved provider actions.
    Never request, print, or store credentials.
    Always store results in `result`.
    """

Do not access _auth_manager from setup() because authentication is attached later in mesh.start(). If application-owned execution code needs to inspect it, use getattr(self, "_auth_manager", None) during task execution. Connected-app calls require a configured and reachable Nexus Gateway; local-vault provider calls use the sandbox's nexus_call boundary instead.


Multi-Agent Workflows

Steps in a workflow run in dependency order. Steps with no depends_on run immediately; steps that declare dependencies wait for those steps to complete.

results = await mesh.workflow("pipeline-001", [
    {"id": "fetch",   "agent": "fetcher",   "task": "Fetch AAPL price data for the last 30 days"},
    {"id": "analyse", "agent": "analyst",   "task": "Analyse the price trend",   "depends_on": ["fetch"]},
    {"id": "report",  "agent": "reporter",  "task": "Write an executive summary", "depends_on": ["analyse"]},
])

Prior step outputs are delivered to each downstream agent automatically. The WorkflowEngine builds the dependency output map from every step listed in depends_on, places it in context["previous_step_results"], and the ContextManager renders it as a clearly labelled section in the LLM's context window before every decision turn. The agent reads what the prior step produced and acts on it without any additional code or system prompt instructions from the developer.

The depends_on declaration is the only thing required. A step that does not declare a dependency does not receive the output of steps it has no declared relationship with, even if those steps happened to finish earlier.

Steps without depends_on that share no dependency chain run concurrently. The WorkflowEngine dispatches them in parallel automatically.

Workflow execution is crash-safe. If the process restarts with the same workflow_id, completed steps are not re-run. Only pending steps resume.


Distributed Mesh Goals

Do not confuse the two goal APIs:

  • await agent.execute_goal(...) runs one AutoAgent's Plan, Execute, Evaluate loop described below.
  • await mesh.execute_goal(...) compiles a source goal into a Redis-backed DAG whose steps are claimed independently by capable peers.

An AutoAgent participates in a distributed goal through its normal execute_task() method. No subclass change is required. Declare provider authority when planning needs to distinguish reads, proposals and effects:

Human goal to a peer-claimed DAG

flowchart LR
    Human["Human objective"] --> Register["Mesh.execute_goal()"]
    Register --> Lease["Temporary planning lease"]
    Lease --> Plan["Obligations + capability DAG"]
    Plan --> Ready["Durable ready steps"]
    Ready --> ClaimA["Eligible peer claims step A"]
    Ready --> ClaimB["Eligible peer claims step B"]
    ClaimA --> Attempts["Immutable attempts + artifacts"]
    ClaimB --> Attempts
    Attempts --> Truth["Current obligation projection"]
    Truth --> Response["Current-revision response"]

The planner chooses capabilities, not privileged agent identities. Any eligible peer may claim a ready step under a renewable lease. The planning node can also execute work if it owns the capability, but it has no permanent coordinator role and no special authority after publishing the DAG.

class PipelineAgent(AutoAgent):
        role = "pipeline"
        capabilities = ["pipeline_inspection"]
        capability_descriptions = {
                "pipeline_inspection": "Inspect and reconcile CRM pipeline state.",
        }
        capability_contracts = {
                "pipeline_inspection": {
                        "effects": ["read", "propose"],
                        "systems": ["hubspot"],
                },
        }

The distributed task context includes the exact source objective, workflow and obligation ledger, durable dependency artifacts, dependency interpretations, current capability/effect/systems, and the shared execution budget. A normal successful AutoAgent result satisfies the obligations covered by its step. Applications that distinguish attempt completion from evidence satisfaction may override execute_task() and add the optional semantic interpretation envelope described in Durable Goal Execution.

When attached to a started Mesh, AutoAgent reasoning receives these peer and workflow tools automatically:

Tool Purpose
list_peers() Read online peers and capabilities
ask_peer(role, question) Request help from a specific peer role
ask_capability(capability, question) Request any peer owning a capability
broadcast_update(message) Notify every peer without creating a request
read_mailbox(limit=10) Read durable unread peer notifications
inspect_workflow() Read this goal, obligations, steps and live statuses
read_workflow_step(step_id) Read one durable step definition and output

The peer decides its own tools and provider actions. ask_capability selects an authority boundary, not an agent implementation. Workflow inspection requires Redis; mailbox reading requires a configured mailbox.

For result status, selective revisions, cancellation and deployment rules, see Durable Goal Execution.


Goal-Oriented Execution

Setting goal_oriented = True switches the agent from a single OODA loop to a Plan, Execute, Evaluate loop. The agent decomposes the goal into steps, executes each through the Kernel, evaluates the outcome, and replans automatically if a step fails.

Plan mode does not mean the agent plans everything. With goal_oriented = True the agent triages each task first: simple, single-answer work routes straight to the Kernel for a direct turn, and only genuinely multi-step work goes through the planner. Plan mode makes the agent planning-capable, not planning-forced.

class ResearchAgent(AutoAgent):
    role = "researcher"
    capabilities = ["research", "analysis"]
    system_prompt = "You are a market researcher. Always store results in `result`."
    goal_oriented = True

The execute_task() interface and mesh.workflow() call are identical. The response gains a goal_execution summary:

step["status"]                   # "success" | "failure" | "hitl"
step["payload"]                  # final synthesised answer
step["goal_execution"]              # planning or direct-Kernel summary

Control the loop ceiling with environment variables:

Variable Default Description
MAX_GOAL_STEPS 30 Hard ceiling on plan steps
MAX_REPLAN_ATTEMPTS 8 Maximum replanning cycles before the goal fails
MAX_PARALLEL_STEPS 3 Concurrency ceiling for independent plan steps

Parallel steps in plans

The planner declares depends_on on each step, listing only the steps whose output it actually needs. Steps with no ordering constraint between them run at the same time, up to MAX_PARALLEL_STEPS. A step never starts before all of its dependencies have finished with a passing result. Plans that declare no depends_on at all run one step at a time, exactly as before.

plan: [
  ("step_01_gather_risks",      []),                          # ┐ run in
  ("step_02_gather_mitigations", []),                          # ┘ parallel
  ("step_03_write_brief", ["step_01_gather_risks",
                           "step_02_gather_mitigations"]),     # waits for both
]

Goal persistence and resume

Goal executions save themselves automatically when blob storage is attached. A snapshot is written after planning, after every completed step, and when the goal ends, under goals/{agent_id}/{goal_id}.json. If the process crashes, you lose at most the step that was running:

execution = await agent.execute_goal("Produce the Q2 analysis")
goal_id = execution.goal_id                      # capture for resume

# After a crash or restart:
execution = await agent.execute_goal(
    "Produce the Q2 analysis",
    resume_goal_id=goal_id,                      # rehydrates plan, facts, history
)

Resume continues from the first step that has not passed yet. If the snapshot is missing or unreadable, the agent logs a warning and starts fresh. Resume never crashes a goal.

For the CustomAgent side of this boundary, where planning is a library you call rather than a mode you enable, see Planning: a library, not a mode.


Human-in-the-Loop Escalation

In goal-oriented mode, if the evaluator's confidence on a step falls below the configured threshold, the Kernel pauses and emits a HITL event. Your application receives this and presents it to a human operator. The operator response is injected as the step result.

.env
HITL_ENABLED=true
HITL_MAX_CONFIDENCE=0.8
HITL_MIN_RISK_SCORE=0.7

Per-agent escalation targets are configured in the agent profile YAML under escalates_to. See the Agent Personas concept page for the full schema. For the full HITL API (request(), wait(), check(), resolve()) see the HITL Escalation guide.


Production Example: Financial Pipeline

The financial_pipeline.py example runs three AutoAgents sequentially (a market data fetcher, an analyst, and a report writer) with Redis crash recovery, UnifiedMemory episodic logging, and blob storage output.

docker compose -f docker-compose.infra.yml up -d
cp .env.example .env   # set GEMINI_API_KEY and REDIS_URL
python examples/financial_pipeline.py

Verify the outputs:

redis-cli hgetall "step_output:financial-daily-001:fetch"
cat blob_storage/reports/financial-daily-001.md
redis-cli xrange ledgers:financial-daily-001 - +

Running with the same workflow_id a second time skips already-completed steps. See the Financial Pipeline example for the full annotated walkthrough.


Troubleshooting

If a task returns a success status with execution_time under 10 milliseconds and a null payload, the generated code failed instantly before doing any real work. This almost always means context was not passed in the task dict, so the generated code raised a NameError on the first line that referenced it. Make sure every step dict includes a "context" key, even if empty.

If code keeps failing after autonomous repairs, improve the system prompt. The LLM generates code from the prompt. Vague prompts produce fragile code: the more explicit you are about available tools, input format, and the exact shape of result, the fewer repairs are needed.