AutoAgent Guide¶
AutoAgent is JarvisCore's autonomous reasoning profile. A minimal agent needs
three class attributes: role, capabilities, and system_prompt. Production
agents can also describe each capability, declare its authorized effects and
systems, validate output, select an execution role, and request Nexus-backed
credentials. The framework then handles LLM selection, Kernel routing,
sandboxed execution, repair, and peer participation. For multi-step goals,
goal_oriented = True enables Plan, Execute, Evaluate with automatic replanning.
Defining an AutoAgent¶
from jarviscore import AutoAgent
class ResearcherAgent(AutoAgent):
role = "researcher"
capabilities = ["research", "synthesis", "web-search"]
system_prompt = """
You are a rigorous research analyst. Prioritise primary sources,
cross-reference findings, and structure outputs as actionable intelligence.
Always store your final output in a variable named `result`.
"""
The framework requires role, capabilities, and system_prompt; construction
raises ValueError when one is absent. That is the smallest valid AutoAgent,
not the complete production contract. Add the optional declarations below when
the Mesh or Kernel needs them to route, authorize, or validate work.
Class Attributes¶
| Attribute | Required | Description |
|---|---|---|
role |
Yes | Slug used for peer discovery, profile loading, and workflow routing |
capabilities |
Yes | Tags for capability-based peer discovery |
system_prompt |
Yes | Base LLM system prompt; framework raises ValueError if absent |
description |
No | One-sentence purpose used by peers for routing decisions |
capability_descriptions |
No | Mapping from capability name to a concrete routing description. Distributed planning uses it instead of guessing from a short tag. |
capability_contracts |
No | Mapping from capability name to authorized effects, provider systems, and optional planner-visible produces artifact descriptions. |
output_schema |
No | Pydantic model class enforced on CoderSubAgent execution output. Validate other role outputs or richer work products in an application subclass. |
default_kernel_role |
No | Preferred fallback role for specialist agents. Built-ins are "researcher", "coder", "communicator", and "browser"; products may register custom roles through an extended Kernel. Leave unset for generalists. |
goal_oriented |
No | Defaults to False; set True for multi-step goal decomposition |
requires_auth |
No | Defaults to False; set True to receive Nexus-backed _auth_manager |
Production Mesh declaration¶
The class remains an ordinary AutoAgent subclass. The additional attributes
make its routing and authority explicit:
from jarviscore import AutoAgent
class RepositoryReviewer(AutoAgent):
role = "repository_reviewer"
description = "Reviews repository changes against correctness evidence."
capabilities = ["code_review"]
capability_descriptions = {
"code_review": "Inspect a change, identify defects, and propose review findings.",
}
capability_contracts = {
"code_review": {
"effects": ["read", "propose"],
"systems": ["github"],
"produces": "ReviewReport(findings, inspected_revision)",
},
}
default_kernel_role = "coder"
requires_auth = True
system_prompt = """
You are a rigorous code reviewer. Ground every finding in repository
evidence and return only review findings supported by the inspected change.
Always store the final output in `result`.
"""
effects may contain read, propose, write, notify, or destructive.
systems names the provider boundaries the capability may use. These contracts
do not grant credentials by themselves: requires_auth controls credential
injection, and provider policy still governs each operation.
Running an AutoAgent¶
Standalone script¶
import asyncio
from jarviscore import Mesh
from agents import ResearcherAgent
async def main():
mesh = Mesh()
mesh.add(ResearcherAgent)
await mesh.start()
results = await mesh.workflow("research-001", [
{"agent": "researcher", "task": "Summarise the state of AI hardware in 2026"}
])
step = results[0]
print(step["status"]) # "success"
print(step["payload"]) # the value stored in `result` by the generated code
await mesh.stop()
asyncio.run(main())
Mesh() takes no mode argument. It auto-detects available infrastructure at start() time. Pass REDIS_URL in your environment and the Mesh will use Redis automatically. Pass it explicitly via the config dict if you need to override:
For P2P between nodes, add p2p_enabled: True and bind_port to the config dict. No mode string required.
FastAPI service¶
from contextlib import asynccontextmanager
from fastapi import FastAPI
from jarviscore import Mesh
from agents import ResearcherAgent
mesh = Mesh()
@asynccontextmanager
async def lifespan(app: FastAPI):
mesh.add(ResearcherAgent)
await mesh.start()
yield
await mesh.stop()
app = FastAPI(lifespan=lifespan)
@app.post("/research")
async def research(request: dict):
results = await mesh.workflow("req-001", [
{"agent": "researcher", "task": request["task"]}
])
step = results[0]
return {"status": step["status"], "output": step.get("payload")}
What workflow() Returns¶
mesh.workflow() returns a list of step result dicts in the original declared
order. Each result carries the agent envelope plus its stable step_id:
| Key | Description |
|---|---|
status |
"success", "failure", "yield", or "hitl" for a paused goal-oriented execution |
output |
Programmatic task result |
payload |
Alias of output on the standard Kernel path |
result_summary |
Guaranteed plain-prose display summary |
error |
None on success; failure or yield explanation otherwise |
tokens |
Input, output, and total token counts |
cost_usd |
Estimated execution cost |
step_id |
Stable workflow step identity |
goal_execution |
Present for goal-oriented planning or a classified direct Kernel turn |
Always check step["status"] == "success" before consuming step["output"].
A failure has error plus a human-readable result_summary; a yield or HITL
result carries continuation details in yield_metadata.
When calling agent.execute_task() directly, the returned envelope always carries result_summary: plain prose suitable for display, never JSON. Failure envelopes put a human-readable error sentence there. Structured data stays in output / payload / goal_execution for programmatic consumers.
Lifecycle Hooks¶
AutoAgent.setup() is called once by the Mesh after instantiation. It initialises the Kernel, LLM client, sandbox, FunctionRegistry, and loads the agent persona from YAML if configured. Override it for one-time setup:
async def setup(self):
await super().setup() # must come first
self.data_client = await MyDataClient.connect()
teardown() is called during Mesh shutdown. Release any resources opened in setup():
You will rarely need to override execute_task(). The Kernel pipeline routes internally. Override it only if you need to enrich the task dict before the Kernel processes it:
async def execute_task(self, task: dict) -> dict:
enriched = {**task, "task": f"{task.get('task', '')}\n\nUser context: {self.user_id}"}
return await super().execute_task(enriched)
System Prompt Best Practices¶
The system prompt is the primary way to shape the code the agent generates.
system_prompt = """
You are a financial data analyst.
Data source: Yahoo Finance API at https://query1.finance.yahoo.com/v8/finance/chart/{ticker}
Parse the JSON response carefully and handle missing keys with .get().
Your output must be stored in a variable named `result` as a dict with keys:
ticker (str)
price (float)
change_pct (float)
analysis (str, 2-3 sentences)
"""
Tell the agent exactly what variable name to use (result), the expected shape of that variable, what APIs or tools are available in the sandbox, and what to do when data is missing or a request fails. Vague prompts produce fragile code.
Kernel Role Routing¶
The Kernel classifies each task and routes it to the appropriate sub-agent. You do not configure this directly. The Kernel infers the right sub-agent from the task. On specialist agents where every task always uses the same sub-agent, skip classification entirely by setting default_kernel_role:
class SlackNotifier(AutoAgent):
role = "notifier"
default_kernel_role = "communicator" # always sends, never codes or searches
The four sub-agents the Kernel routes to are the CoderSubAgent for tasks that require writing and running Python, the ResearcherSubAgent for tasks that require web search and synthesis, the CommunicatorSubAgent for tasks that require formatting and delivering output, and the BrowserSubAgent for web navigation tasks when BROWSER_ENABLED=true is set.
Analysis, not code: the single_response contract¶
The answer must contain visible text and must not carry an explicit incomplete,
filtered, refused, tool-call or unrecognized terminal reason. Otherwise the result
uses status="failure", retaining any partial output/payload, finish_reason,
provider_metadata, provider/model, tokens and cost for diagnosis. It never retries
or reports a new warning status. Nonempty responses from older/custom clients with
no finish reason remain compatible; token counts alone never establish truncation.
To set an output budget for this turn only, include a positive integer
max_output_tokens in the execution contract, for example:
{"execution_shape": "single_response", "max_output_tokens": 8192}.
This is forwarded as generate(max_tokens=8192). Invalid values fail before an
LLM call. Omitting it retains the existing configured provider budget; model
limits still apply. A reasoning model may spend part of that budget on internal
reasoning rather than visible text.
Many agent tasks need exactly one LLM completion: render the system prompt, ask the question, return the answer. No planner, no routing, no code generation. Declare this shape per task with an execution contract:
results = await mesh.workflow("analysis-001", [{
"agent": "market_analyst",
"task": "Analyse EURUSD H1 and return your thesis as JSON.",
"context": {"execution_contract": {"execution_shape": "single_response"}},
}])
With single_response declared, execute_task() runs one completion against the agent's system prompt (including its persona profile, if one is loaded) and returns the standard result envelope with token and cost telemetry. The Kernel pipeline is skipped entirely, so an analysis prompt never reaches the Coder sub-agent and never produces TOOL/DONE protocol errors.
Use this for tasks where the whole job is the answer: analysis, classification, extraction, drafting. When the task needs tools, search, or several turns of reasoning, drop the contract and let the Kernel route it.
Coder Sandbox¶
When the Kernel routes a task to the CoderSubAgent, code executes inside CoderSandbox: a deliberately file-capable execution environment that is distinct from the SandboxExecutor used by other sub-agents.
SandboxExecutor (used by ResearcherSubAgent and CommunicatorSubAgent) blocks open(), subprocess, and filesystem access. CoderSandbox intentionally grants those capabilities, scoped to a controlled workspace directory.
Inside every execution, generated code receives these names in its namespace:
| Name | Type | Description |
|---|---|---|
workspace |
Path |
Project root directory. All file reads and writes are expected here. |
output_dir |
Path |
workspace/output/: the preferred write location for produced files. |
blob_path(name) |
Path |
Shorthand for output_dir / name. Creates parent directories. |
bash(cmd) |
BashExecutor |
Run allowed shell commands. Returns {success, stdout, stderr, returncode}. |
git |
GitHelper |
High-level git: checkout_branch, add_all, commit, push, describe_pr. |
nexus_call |
async fn | Make authenticated HTTP calls to any provider via Nexus. Never exposes credentials. |
json, os, re, Path, datetime, math, shutil, uuid |
stdlib | Pre-imported for convenience. |
The Coder can import any installed package from within the sandbox. The security boundary is the bash allow-list, not Python builtins.
The bash allow-list covers: git, pip, pip3, cp, mv, mkdir, rm, ls, cat, echo, touch, find, grep, sed, awk, pandoc, npm, npx, node, python, python3, curl. Hard-blocked regardless: sudo, rm -rf /, eval, curl | bash, and pipe-to-shell patterns.
The CoderResult output shape: what gets written to result by generated code and surfaced in step["payload"]:
result = {
"success": bool, # did the task complete?
"files_created": [str], # absolute paths of files written
"files_modified": [str], # absolute paths of files changed
"git_branch": str | None, # branch name if git ops were performed
"stdout": str, # captured print() output
"data": Any, # any structured data to return
"error": str | None, # error message if success=False
}
Configure the workspace root and timeouts:
# Workspace directory: defaults to the process working directory
CODER_WORKSPACE=/path/to/your/project
# Sandbox execution timeout in seconds (default: 300)
SANDBOX_TIMEOUT=300
When writing system prompts for agents that route to the Coder, always tell the agent what files it should write and where. The Coder will use output_dir by default if you do not specify a location.
Execution Budgets¶
Every sub-agent dispatch runs against an ExecutionLease: a token, turn, and wall-clock budget specific to the sub-agent role. When the lease is exhausted, the Kernel stops the sub-agent and returns whatever has been produced so far.
| Role | Thinking tokens | Action tokens | Total tokens | Wall clock | Turn fuse |
|---|---|---|---|---|---|
coder |
132,000 | 108,000 | 240,000 | 4 min | 32 turns |
researcher |
180,000 | 60,000 | 240,000 | 4 min | 36 turns |
communicator |
72,000 | 48,000 | 120,000 | 4 min | 18 turns |
browser |
60,000 | 60,000 | 120,000 | 5 min | 28 turns |
The turn fuse is an emergency hard-stop. If a sub-agent reaches 32 turns without completing, the Kernel terminates it regardless of token budget. This prevents runaway loops on adversarial or ambiguous tasks.
You do not configure lease budgets per-agent: they are role-level defaults. If a specific task consistently exhausts the budget, the right fix is usually to narrow the task scope (break it into smaller steps in the workflow DAG), not to increase the budget.
Token consumption appears in step["tokens"] on every AutoAgent workflow result:
results = await mesh.workflow("report", [
{"id": "analyse", "agent": "analyst", "task": "..."}
])
step = results[0]
print(step["tokens"])
# {"input": 14200, "output": 3100, "total": 17300}
Model Routing¶
JarvisCore routes each sub-agent dispatch to a capability tier (nano, standard, or heavy) based on the role and an optional per-step hint. Pass complexity in a workflow step to override the default for that dispatch:
results = await mesh.workflow("task-001", [
{"agent": "analyst", "task": "Summarise this paragraph.", "complexity": "nano"},
{"agent": "analyst", "task": "Produce a competitive analysis.", "complexity": "heavy"},
])
When complexity is omitted, the role's built-in default applies: communicator defaults to nano, researcher and browser default to standard. See Model Routing for the full tier configuration, fallback chain, and provider compatibility reference.
Infrastructure Injection¶
The Mesh injects stores, mailbox, and HITL before setup() runs. Peer clients
and optional authentication are attached later in mesh.start() and are
available before the Mesh accepts work.
| Attribute | Available when |
|---|---|
self._redis_store |
REDIS_URL is set |
self._blob_storage |
Always: falls back to local filesystem |
self.mailbox |
Always: in-memory by default; Redis-backed when configured |
self.hitl |
When the framework HITL component is available; Redis-enhanced when configured |
self.peers |
After mesh.start() attaches local or distributed peer transport |
self._auth_manager |
After setup(), when requires_auth = True and authentication is configured |
from jarviscore.memory import UnifiedMemory
async def setup(self):
await super().setup()
self.memory = UnifiedMemory(
workflow_id="wf-001",
step_id=self.role,
agent_id=self.role,
redis_store=self._redis_store,
blob_storage=self._blob_storage,
)
Agents running inside a workflow write their step output to step_output:{workflow_id}:{step_id} in Redis. Read a prior step's output from another agent using RedisMemoryAccessor:
from jarviscore.memory import RedisMemoryAccessor
accessor = RedisMemoryAccessor(self._redis_store, workflow_id="wf-001")
raw = accessor.get("fetch")
prior = raw.get("output", raw) if isinstance(raw, dict) else {}
Checkpointing¶
Checkpointing happens automatically. After every OODA loop turn, the Kernel calls UnifiedMemory.save_checkpoint() which writes the current KernelState (all accumulated findings, tool history, and reasoning progress) to Redis under checkpoint:{workflow_id}:{step_id}. If the process crashes and the workflow restarts, the WorkflowEngine detects the existing checkpoint and resumes from that turn rather than restarting from scratch.
You do not call save_checkpoint() or load_checkpoint() in application code. The framework manages this transparently when REDIS_URL is configured. Without Redis, there is no persistence and no crash recovery: the workflow starts from the beginning on failure.
Nexus Auth: requires_auth¶
Set requires_auth = True on agents that call connected third-party services.
When connected-app authentication is configured, the Mesh creates an
AuthenticationManager and injects it as self._auth_manager after setup().
The Kernel exposes nexus_call inside the sandbox; generated code never receives
the manager or raw credentials.
class GitHubAgent(AutoAgent):
role = "github_agent"
capabilities = ["github", "code-review"]
requires_auth = True
system_prompt = """
Use the available Nexus-backed GitHub tools for approved provider actions.
Never request, print, or store credentials.
Always store results in `result`.
"""
Do not access _auth_manager from setup() because authentication is attached
later in mesh.start(). If application-owned execution code needs to inspect it,
use getattr(self, "_auth_manager", None) during task execution. Connected-app
calls require a configured and reachable Nexus Gateway; local-vault provider
calls use the sandbox's nexus_call boundary instead.
Multi-Agent Workflows¶
Steps in a workflow run in dependency order. Steps with no depends_on run immediately; steps that declare dependencies wait for those steps to complete.
results = await mesh.workflow("pipeline-001", [
{"id": "fetch", "agent": "fetcher", "task": "Fetch AAPL price data for the last 30 days"},
{"id": "analyse", "agent": "analyst", "task": "Analyse the price trend", "depends_on": ["fetch"]},
{"id": "report", "agent": "reporter", "task": "Write an executive summary", "depends_on": ["analyse"]},
])
Prior step outputs are delivered to each downstream agent automatically. The WorkflowEngine builds the dependency output map from every step listed in depends_on, places it in context["previous_step_results"], and the ContextManager renders it as a clearly labelled section in the LLM's context window before every decision turn. The agent reads what the prior step produced and acts on it without any additional code or system prompt instructions from the developer.
The depends_on declaration is the only thing required. A step that does not declare a dependency does not receive the output of steps it has no declared relationship with, even if those steps happened to finish earlier.
Steps without depends_on that share no dependency chain run concurrently. The WorkflowEngine dispatches them in parallel automatically.
Workflow execution is crash-safe. If the process restarts with the same workflow_id, completed steps are not re-run. Only pending steps resume.
Distributed Mesh Goals¶
Do not confuse the two goal APIs:
await agent.execute_goal(...)runs one AutoAgent's Plan, Execute, Evaluate loop described below.await mesh.execute_goal(...)compiles a source goal into a Redis-backed DAG whose steps are claimed independently by capable peers.
An AutoAgent participates in a distributed goal through its normal
execute_task() method. No subclass change is required. Declare provider
authority when planning needs to distinguish reads, proposals and effects:
Human goal to a peer-claimed DAG¶
flowchart LR
Human["Human objective"] --> Register["Mesh.execute_goal()"]
Register --> Lease["Temporary planning lease"]
Lease --> Plan["Obligations + capability DAG"]
Plan --> Ready["Durable ready steps"]
Ready --> ClaimA["Eligible peer claims step A"]
Ready --> ClaimB["Eligible peer claims step B"]
ClaimA --> Attempts["Immutable attempts + artifacts"]
ClaimB --> Attempts
Attempts --> Truth["Current obligation projection"]
Truth --> Response["Current-revision response"]
The planner chooses capabilities, not privileged agent identities. Any eligible peer may claim a ready step under a renewable lease. The planning node can also execute work if it owns the capability, but it has no permanent coordinator role and no special authority after publishing the DAG.
class PipelineAgent(AutoAgent):
role = "pipeline"
capabilities = ["pipeline_inspection"]
capability_descriptions = {
"pipeline_inspection": "Inspect and reconcile CRM pipeline state.",
}
capability_contracts = {
"pipeline_inspection": {
"effects": ["read", "propose"],
"systems": ["hubspot"],
},
}
The distributed task context includes the exact source objective, workflow and
obligation ledger, durable dependency artifacts, dependency interpretations,
current capability/effect/systems, and the shared execution budget. A normal
successful AutoAgent result satisfies the obligations covered by its step.
Applications that distinguish attempt completion from evidence satisfaction may
override execute_task() and add the optional semantic interpretation
envelope described in Durable Goal Execution.
When attached to a started Mesh, AutoAgent reasoning receives these peer and workflow tools automatically:
| Tool | Purpose |
|---|---|
list_peers() |
Read online peers and capabilities |
ask_peer(role, question) |
Request help from a specific peer role |
ask_capability(capability, question) |
Request any peer owning a capability |
broadcast_update(message) |
Notify every peer without creating a request |
read_mailbox(limit=10) |
Read durable unread peer notifications |
inspect_workflow() |
Read this goal, obligations, steps and live statuses |
read_workflow_step(step_id) |
Read one durable step definition and output |
The peer decides its own tools and provider actions. ask_capability selects an
authority boundary, not an agent implementation. Workflow inspection requires
Redis; mailbox reading requires a configured mailbox.
For result status, selective revisions, cancellation and deployment rules, see Durable Goal Execution.
Goal-Oriented Execution¶
Setting goal_oriented = True switches the agent from a single OODA loop to a Plan, Execute, Evaluate loop. The agent decomposes the goal into steps, executes each through the Kernel, evaluates the outcome, and replans automatically if a step fails.
Plan mode does not mean the agent plans everything. With goal_oriented = True
the agent triages each task first: simple, single-answer work routes straight to
the Kernel for a direct turn, and only genuinely multi-step work goes through the
planner. Plan mode makes the agent planning-capable, not planning-forced.
class ResearchAgent(AutoAgent):
role = "researcher"
capabilities = ["research", "analysis"]
system_prompt = "You are a market researcher. Always store results in `result`."
goal_oriented = True
The execute_task() interface and mesh.workflow() call are identical. The response gains a goal_execution summary:
step["status"] # "success" | "failure" | "hitl"
step["payload"] # final synthesised answer
step["goal_execution"] # planning or direct-Kernel summary
Control the loop ceiling with environment variables:
| Variable | Default | Description |
|---|---|---|
MAX_GOAL_STEPS |
30 |
Hard ceiling on plan steps |
MAX_REPLAN_ATTEMPTS |
8 |
Maximum replanning cycles before the goal fails |
MAX_PARALLEL_STEPS |
3 |
Concurrency ceiling for independent plan steps |
Parallel steps in plans¶
The planner declares depends_on on each step, listing only the steps whose output it actually needs. Steps with no ordering constraint between them run at the same time, up to MAX_PARALLEL_STEPS. A step never starts before all of its dependencies have finished with a passing result. Plans that declare no depends_on at all run one step at a time, exactly as before.
plan: [
("step_01_gather_risks", []), # ┐ run in
("step_02_gather_mitigations", []), # ┘ parallel
("step_03_write_brief", ["step_01_gather_risks",
"step_02_gather_mitigations"]), # waits for both
]
Goal persistence and resume¶
Goal executions save themselves automatically when blob storage is attached. A snapshot is written after planning, after every completed step, and when the goal ends, under goals/{agent_id}/{goal_id}.json. If the process crashes, you lose at most the step that was running:
execution = await agent.execute_goal("Produce the Q2 analysis")
goal_id = execution.goal_id # capture for resume
# After a crash or restart:
execution = await agent.execute_goal(
"Produce the Q2 analysis",
resume_goal_id=goal_id, # rehydrates plan, facts, history
)
Resume continues from the first step that has not passed yet. If the snapshot is missing or unreadable, the agent logs a warning and starts fresh. Resume never crashes a goal.
For the CustomAgent side of this boundary, where planning is a library you call rather than a mode you enable, see Planning: a library, not a mode.
Human-in-the-Loop Escalation¶
In goal-oriented mode, if the evaluator's confidence on a step falls below the configured threshold, the Kernel pauses and emits a HITL event. Your application receives this and presents it to a human operator. The operator response is injected as the step result.
Per-agent escalation targets are configured in the agent profile YAML under escalates_to. See the Agent Personas concept page for the full schema. For the full HITL API (request(), wait(), check(), resolve()) see the HITL Escalation guide.
Production Example: Financial Pipeline¶
The financial_pipeline.py example runs three AutoAgents sequentially (a market data fetcher, an analyst, and a report writer) with Redis crash recovery, UnifiedMemory episodic logging, and blob storage output.
docker compose -f docker-compose.infra.yml up -d
cp .env.example .env # set GEMINI_API_KEY and REDIS_URL
python examples/financial_pipeline.py
Verify the outputs:
redis-cli hgetall "step_output:financial-daily-001:fetch"
cat blob_storage/reports/financial-daily-001.md
redis-cli xrange ledgers:financial-daily-001 - +
Running with the same workflow_id a second time skips already-completed steps. See the Financial Pipeline example for the full annotated walkthrough.
Troubleshooting¶
If a task returns a success status with execution_time under 10 milliseconds and a null payload, the generated code failed instantly before doing any real work. This almost always means context was not passed in the task dict, so the generated code raised a NameError on the first line that referenced it. Make sure every step dict includes a "context" key, even if empty.
If code keeps failing after autonomous repairs, improve the system prompt. The LLM generates code from the prompt. Vague prompts produce fragile code: the more explicit you are about available tools, input format, and the exact shape of result, the fewer repairs are needed.