Admit agent runs only when bounded execution and queue capacity remain. Limit queued and running work separately, track oldest-work age alongside depth, and define what happens at overflow: reject, shed stale or low-value work, or replace superseded work. Model-call throttling alone does not prevent memory growth, starvation, or unactionable queue latency.
Treat caller cancellation and execution timeout as separate signals, propagate both through every model and tool call, and assign each step a deadline within one total run budget. Reserve time for cleanup and final reporting, so a disconnected client or late tool cannot leave paid work running or erase the run’s outcome.
Keep a fan-out agent inside provider rate limits by bounding concurrency with a worker pool, tracking both the request-per-minute and token-per-minute axes, honouring the Retry-After header on every 429, and adding randomised jitter so retrying workers do not resynchronise into the burst that triggered the limit.
Set retention at the field and store level, not once per agent run. Conversation history, operational traces, and tool payloads serve different purposes; deletion must cascade into embeddings, caches, and evaluation examples. Keep each field only for a documented legal or operational purpose, with an owner, expiry trigger, and verifiable deletion path.
Bound agent disk access outside the process it controls. Restrict reads because repository data enters model context, treat writes to build and CI files as deferred execution, require approval for deletion, and combine filesystem policy with network egress controls so readable secrets cannot become an outbound payload.
For a single agent, a plain tool-calling loop is the default: under 200 lines the team understands, no scaffolding prompt, and stack traces through code you wrote. A framework earns its place when you need multi-agent orchestration, durable resumption or human approval gates, or run enough agents that retries, tracing and persistence are worth writing once.
Handing work between agents is reliable only when the boundary carries a compact task contract: objective, constraints, current state, idempotency key, correlation identifier, and failure policy. Transcripts are poor handoff records. The receiver must be able to retry safely, trace the run across systems, and recover when no receiver responds.
Give every agent a distinct identity and credentials. Human credentials misattribute automated actions, broaden the agent’s permissions and couple revocation to the operator’s access. Preserve the agent identity across downstream calls; otherwise attribution ends at a shared token. Accept the added work of provisioning, rotation and retirement.
Give an agent URL transformation rules and a vendor skill pack for code generation; add an application-owned signed upload endpoint when it must ingest files. Use an MCP server only when the task must inspect or change a live environment. Begin without tools, then add the blocked capability.
An agent should persist working state for the current task in a structure it writes deliberately, keep durable knowledge in a store with an eviction policy, and treat conversation history as disposable. Anything persisted must be human-readable and provenance-tagged, because a later run trusts it. The test: a run resumes on another machine from the store alone.
An agent can provision a working Cloudinary environment mid-session by running one unauthenticated npx command, which writes credentials to a local env file and prints a claim URL. The environment expires after 24 hours unless a person claims it, and claiming keeps the same credentials, so nothing written against it needs rewriting.
To reconstruct an agent run, record one trace per task with a span per model or tool call, log every tool call's arguments alongside its result, store the prompt actually sent rather than the template, and instrument to the OpenTelemetry generative-AI semantic conventions. Replaying the model is a new run; the tool sequence is what replays exactly.
Plan only far enough to expose assumptions, dependencies, irreversible actions, and review points. Let the agent act directly on cheap, reversible work, preserve an explicit plan for human objection, split only independent subtasks, and re-plan when evidence shows the current route has failed rather than after every successful step.
A permission model for an agent decides which operations run unattended and which pause for a human. The workable axis is reversibility, not trust: reads and additive writes run freely; deletes, publishes and payments prompt. Enforce it by allowlisting tools per session, scoping and time-bounding credentials, and keeping an escape hatch a human can reach mid-run.
Put a separate circuit breaker around each meaningful tool dependency. Count only availability-related failures, open the breaker when a recent threshold is crossed, reject new calls during recovery, and let only a few probes through half-open. Keep retries bounded behind the breaker, with backoff and jitter, so they cannot recreate the outage.
Treat every agent side effect as a durable forward-and-undo record: save the result, compensation inputs, and authorization before continuing. Declare the point of no return, make everything after it retryable or manually resolvable, and run compensation as an idempotent, checkpointed workflow with explicit escalation when an undo cannot finish.
Keep long agent runs useful by shrinking the always-loaded prompt and tool surface, admitting only relevant results, and moving editable conclusions into durable state. Compact history only at controlled checkpoints, treat prompt caching as a cost measure, and detect context failure with instruction-following evaluations rather than waiting for an overflow error.
Use a business operation ID to join retries and replacement runs, and a trace ID to join telemetry within each distributed execution. Propagate both explicitly through model, tool, and queue calls. Keep every identifier opaque and free of customer data so logs and headers remain safe.
Cost attribution for agents means tying every model call back to the unit of work that triggered it. It is a tracing problem: propagate a task identifier into each call, record it as a span with token counts and cost, and account for cached input tokens separately from uncached ones.
A dead-letter queue stops repeatedly failing agent jobs from exhausting the normal retry path. Each record must preserve the original payload, the complete failure history, and a correlation ID. Operators should replay only after changing the payload or repairing the failed dependency; otherwise the job is expected to fail again.
Durable execution persists the position of a running workflow so a crash resumes from the last completed step rather than restarting the sequence. It is bought with a determinism constraint on workflow code, configured per activity for retries and timeouts, and is worth adopting only when losing a run midway is unacceptable.
Filesystem and process isolation alone does not stop an agent from exfiltrating data, because one outbound request is enough. Default-deny egress with a host allowlist, DNS included, enforced outside the sandboxed process, plus short-lived narrowly scoped credentials, fixes more risk per hour of work than a stronger isolation tier.
A resource an agent creates on a user's behalf expires unless a person claims it. Cloudinary's provisioned product environment deletes itself 24 hours after creation, and claiming requires an email address plus a confirmation from that mailbox, so an unattended run leaves lost work rather than orphaned assets and silent billing.
Use fallback only for retryable transport failures, within the run's deadline. Refusals, invalid tool requests, and contract violations stay on the failure path. A fallback is eligible only when it preserves tool and output contracts and can hold the context required to resume safely.
A golden set gives a stable, auditable score but costs human effort to build and maintain; a model judge scales to any volume while adding a second non-deterministic system that needs its own validation. Use a small golden set to calibrate the judge, then let the judge cover the rest.
An approval gate is a pause-and-resume problem before it is a UX one: persist the run so it can wait, batch the questions so a reviewer reads them, show the diff rather than the intent, expire unanswered requests to deny by default with a visible surface, and store the approval record with the changed artefact.
Protect every side-effecting agent action with one idempotency key tied to the intended operation, persist that key across retries, reject changed payloads, and cache the completed result for at least the full retry window. A retry then returns the original outcome instead of repeating the effect.
A claimable Cloudinary environment locks media delivery to the public IP that ran the provisioning command. Uploads and transformations still succeed, so the failure surfaces only when a browser on another machine, or a preview deployment, requests the asset. Pass extra addresses at provisioning time, or claim the environment to remove the restriction.
An operational kill switch is an independent control path that stops an autonomous agent from starting or continuing unsafe work. Define measurable thresholds, decision owners, and switch scope before deployment; stop admissions, revoke tool credentials, contain active runs, preserve evidence, maintain a non-agent fallback, and require explicit evidence before redeployment.
Metered vendor billing charges a single unit against several different resources at once. With Cloudinary, one credit buys 1,000 transformations, 1 GB of storage, or 1 GB of delivered bandwidth. Agent-driven work exhausts whichever axis it touches, and transformation and bandwidth totals keep counting for thirty rolling days.
Set one retry budget across the entire agent run: cap total attempts, elapsed time, and side-effecting calls. Retry only transient failures on idempotent operations or calls protected by an idempotency key. When any limit is spent, return a typed failure or escalate to a human; never let the model start another loop.
Rotate agent secrets by keeping old and new credentials valid during a measured overlap, switching agents to resolve a stable secret reference at each call, and revoking the old credential only after logs show it is no longer used. If dual validity is impossible, drain every run before replacement.
Treat every page as untrusted input, enforce origin and download policy outside the model before launch, separate preparation from side effects, and require a human confirmation immediately before purchases, messages, deletions, or credential entry. Validate the browser state around each action and stop when the observed state leaves the approved task.
Model-generated code has to run somewhere, and the choice of isolation sets your blast radius. This page compares containers, microVMs, WebAssembly, and hosted sandboxes on startup latency, filesystem and network scope, credential exposure, and operational cost, then states which one each workload profile actually needs.
Run agent-requested code in an isolated environment with no ambient credentials, a minimal filesystem, default-deny network access, and task-scoped capabilities. Enforce separate CPU, memory, wall-time, process-count, output-size, storage, and network limits. Treat the process boundary and syscall filtering as layers, then test every denied path and cleanup action.
Long-running hosts suit agents better on most axes: they bill for machines rather than wall-clock idle, impose no turn deadline, keep model connections and state warm across turns, and absorb bursts without cold starts. Serverless wins for short, bounded, spiky tool calls where per-account concurrency throttling is acceptable.
Use schema-constrained output when software consumes an agent’s result, free text when a person consumes an open-ended answer, and a hybrid when both do. Schemas make malformed results detectable at the boundary; strict rejection preserves that signal, while an unparsed reasoning field retains useful explanation without becoming a hidden dependency.
Keep one agent when the task depends on shared context or produces output the parent cannot verify. Delegate search and summarisation when exploration is large but the return is small, source-backed, and checkable. Limit parallel workers to the provider’s available rate budget.
CI for a non-deterministic system works when you stop asserting on exact output. Assert on properties that survive rephrasing, run each input several times, and gate on a pass rate against a tolerance the product owner sets. Temperature zero narrows variance but never removes it, so design for spread.
Use head sampling when predictable overhead matters and rare outcomes are not the selection target. Use tail sampling when slow, failed, or costly traces must survive, accepting complete-trace buffering. In either design, propagate the same sampling decision through every service so retained traces do not arrive with missing downstream spans.
Cloudinary's agent account-creation endpoint accepts an unauthenticated POST because an agent setting up a new user has no credentials to send. It returns account details, product environment credentials and a machine-readable next-steps block, but those credentials stay inert until the human recipient verifies the account by email and sets a password.
Validate every tool result against a declared success-or-failure schema before it reaches model context. Reject malformed and oversized values as typed failures, never success-like prose. Apply truncation only after the original result passes validation, then shape the validated data to the context budget and record the boundary decision.
A freshly issued credential belongs in a local environment file that version control already ignores, never in the conversation, which is copied into logs and evaluation sets by default, and never in a client-side bundle. Display-time redaction is cosmetic once the value is on disk, so find the vendor's rotation call before you need it.