A provenance record should identify the exact artifact by digest, capture the agent or builder, source revision, tool versions, resolved inputs, policy decisions, and run identifier, then be verified against consumer expectations before use. Recording metadata without checking the bytes and trusted production conditions does not reduce risk.
For production, default to a team-operated server: it makes runs reproducible, observable, and able to use durable state. Keep agents on developer machines only for short, human-owned work. Choose a managed service for sporadic runs when its isolation boundary is acceptable; reject it when sensitive tools or data cannot cross that boundary.
Release an agent change only when the candidate clears aggregate and safety-critical slice gates with uncertainty accounted for, while staying inside fixed cost, latency, turn, and tool-call budgets. Small samples should trigger more evaluation or review, not precise pass/fail claims that one example can reverse.
Treat persisted agent state as a versioned database: expand the schema first, keep new code able to read the previous version, migrate state with idempotent steps, and remove compatibility only after old instances are gone. Test fresh, upgraded, interrupted, and mixed-version paths before making the new shape mandatory.
Version tool schemas as model-facing contracts, not backend implementation details. Requiring a new argument breaks old calls; renaming a tool can alter both dispatch and model selection. Deploy incompatible schemas in parallel under distinct names or explicit versions, and retain per-version traces until the old contract is retired.
Treat every WebSocket reconnect as a new transport, then explicitly resume the existing agent session. Send a stable session ID and the last acknowledged event sequence, replay only later events, and advance the acknowledgement after rendering. That closes disconnect gaps without showing the same agent output twice.
Build agent evaluation datasets from real task distributions, record each request with its initial state and available tools, admit privacy-reviewed production failures, label acceptable outcomes and forbidden actions, remove leakage, and reserve a hidden holdout before prompt tuning. The finished suite should replay cases deterministically and reveal regressions on work that matters.
Canary an agent by releasing a single versioned bundle of model, prompt, tools, and policy to a sticky user cohort. Compare that cohort with the control on predeclared quality, safety, cost, and reliability gates. Roll back by stopping new assignments before draining or safely cancelling active side-effecting runs.
Attach every feedback event to an immutable run record containing the prompt version, model version, tool trace, and final state. Ask separately whether the outcome was correct and whether it was presented well. When it was wrong, capture the expected action or answer and promote that correction into an evaluation target.
Pick a plain loop when prompt control and a portable tool layer matter most. Pick a framework when its tested recovery and tracing remove work you would otherwise build. Pick a managed harness only when its operational help outweighs reduced prompt control and an upgrade cadence you cannot fully govern.
For a twenty-engineer team, buy per-seat only when usage is even across the team; buy usage-metered when a few engineers drive most calls, which is more common. Screen SSO, audit logs, residency and retention first, run a fixed-task trial that measures completion and rework, and set a review date.
Pick cron when bounding how often an agent runs matters more than immediate detection. Pick event-driven triggers when work must begin as soon as a source changes, but make processing duplicate-safe and able to absorb bursts. For important work, use events for speed and a reconciliation cron to repair gaps.
Pin every executable layer: the harness, tool adapters, transitive dependency tree, installer, container base, and external tool protocol versions. Commit the lockfile, but verify registries, builders, and artifact provenance separately. Move every update through human review and tests; never let the agent rewrite the dependencies that define its own execution.
Enforce an agent's spend ceiling in the code path that issues model calls, not in a policy document. Set the ceiling per task, convert token counts to money using each token class's price, define what happens on breach — stop, downgrade, or queue — and count retries against the same ledger, because a restart rebills the work already done.
Mid-tier models match frontier ones on classification, extraction, and routing, where outputs are small and checkable. Frontier models still lead on open-ended reasoning, long-horizon planning, and code that must run first time. Choose per task class, prove it with an eval suite, and set a review date, because prices and capabilities move monthly.
An IDE extension and an MCP server for the same vendor cover different callers. The extension gives a developer a visual surface to browse, search and upload assets; the server exposes those same operations to a model as callable tools that cost context. Only the server's configuration is committable, which makes it the team-level decision.
Measuring a coding assistant's impact means tracking delivery outcomes at the team level rather than usage counts. Lines accepted measures adoption, not value. Vendor productivity percentages come from studies that rarely resemble your team. DORA's four keys measure the system, stay safe to publish, and require a baseline captured before rollout.
Model routing selects which model handles each request, usually to cut cost or latency. The router itself is a component: it adds a decision before every call, with its own latency and failure mode. Heuristic routing on input length or declared task type is cheapest; every model added to the table widens the evaluation surface.
Multi-tenant agent runtimes need tenant identity enforced at every boundary: storage keys, cache keys, queues, logs, and each tool call. Per-tenant state partitions and quotas reduce accidental sharing and noisy neighbors, but sensitive workloads still require explicit storage and process isolation, backed by deny-by-default authorization.
Read shared agent state together with its version, compute the next state from that snapshot, and make the write conditional on the same version. If another writer wins, re-read and recompute before retrying. Keep replayed state updates separate from external effects unless those effects are idempotent and reconcilable.
Prompt caching pays off when repeated agent turns preserve the same long prefix. Put stable tool definitions and standing instructions before changing conversation data, treat any early edit as invalidating later reuse, and measure cached tokens against total input tokens so an attractive hit rate does not hide small absolute savings.
Redact agent traces before export with an explicit allowlist. Keep correlation identifiers, tool names, timing, status, and redacted error classes; discard credentials, API keys, personal data, and full tool payloads. Keep trace baggage equally sparse because it propagates across service boundaries and may be logged downstream.
Store every delayed task in durable scheduler state outside the agent process, attach an immutable execution identity, and claim that identity atomically before side effects. For recurrence, choose catch-up, skip, or coalesce explicitly. Use cron for fixed calendar times and a durable workflow when the scheduled job has multiple steps or waits.
Treat streamed text as a transient draft, route tool events through a separate path, and commit the assistant turn only after an explicit completion event. If the connection ends first, keep the turn incomplete and show either a terminal failure or a resume action tied to the last processed event.
A vendor-published skill pack is instruction text a team installs once into a repository so every assistant answers from the vendor's documented patterns instead of recall. Cloudinary's pack installs with one npx command, offers four selectable skills, and is versioned by the vendor, so assistant behaviour changes without a review step.
Version the exact assembled instruction set, including preambles and retrieved policy, then log that prompt version with the requested and returned model on every run. Gate each immutable prompt release with fixed evaluations, route it through a stable canary cohort, and keep the previous release selectable for immediate rollback.
Authenticate each webhook, use the provider’s event identifier as a unique deduplication key, persist the accepted payload, and hand it to a durable queue before returning a success response. Run the agent only from the queue consumer, where receiver-controlled retries and idempotent processing can absorb duplicate delivery and slow model calls.