Development Choices
Photo of Gregory Mostizky

Gregory Mostizky

Software Engineer

Gregory is a senior engineer with 26 years of experience building at the leading edge of the industry, with roles spanning IBM, Ooyala, Facebook, DoorDash, and now OpenAI. His work has consistently landed in the space before a category exists — early video infrastructure, systems at hyperscale, and the messy, high-ambiguity problems that eventually turn into products.

He is currently building the next generation of AI systems at OpenAI.

Articles by Gregory Mostizky

Backpressure and Admission for Agent Workloads

Admit agent runs only when bounded execution and queue capacity remain. Limit queued and running work separately, track oldest-work age alongside depth, and define what happens at overflow: reject, shed stale or low-value work, or replace superseded work. Model-call throttling alone does not prevent memory growth, starvation, or unactionable queue latency.

Set timeout budgets for AI agent runs

Treat caller cancellation and execution timeout as separate signals, propagate both through every model and tool call, and assign each step a deadline within one total run budget. Reserve time for cleanup and final reporting, so a disconnected client or late tool cannot leave paid work running or erase the run’s outcome.

Running Agents Under Provider Rate Limits

Keep a fan-out agent inside provider rate limits by bounding concurrency with a worker pool, tracking both the request-per-minute and token-per-minute axes, honouring the Retry-After header on every 429, and adding randomised jitter so retrying workers do not resynchronise into the burst that triggered the limit.

Retention Boundaries for Agent Data

Set retention at the field and store level, not once per agent run. Conversation history, operational traces, and tool payloads serve different purposes; deletion must cascade into embeddings, caches, and evaluation examples. Keep each field only for a documented legal or operational purpose, with an owner, expiry trigger, and verifiable deletion path.

Bounding an Agent’s Filesystem Access

Bound agent disk access outside the process it controls. Restrict reads because repository data enters model context, treat writes to build and CI files as deferred execution, require approval for deletion, and combine filesystem policy with network egress controls so readable secrets cannot become an outbound payload.

Agent frameworks vs a plain tool-calling loop

For a single agent, a plain tool-calling loop is the default: under 200 lines the team understands, no scaffolding prompt, and stack traces through code you wrote. A framework earns its place when you need multi-agent orchestration, durable resumption or human approval gates, or run enough agents that retries, tracing and persistence are worth writing once.

Agent Handoffs Across System Boundaries

Handing work between agents is reliable only when the boundary carries a compact task contract: objective, constraints, current state, idempotency key, correlation identifier, and failure policy. Transcripts are poor handoff records. The receiver must be able to retry safely, trace the run across systems, and recover when no receiver responds.

Give Every Agent Its Own Identity

Give every agent a distinct identity and credentials. Human credentials misattribute automated actions, broaden the agent’s permissions and couple revocation to the operator’s access. Preserve the agent identity across downstream calls; otherwise attribution ends at a shared token. Accept the added work of provisioning, rotation and retirement.

Media capability without an MCP server

Give an agent URL transformation rules and a vendor skill pack for code generation; add an application-owned signed upload endpoint when it must ingest files. Use an MCP server only when the task must inspect or change a live environment. Begin without tools, then add the blocked capability.

What an Agent Should Persist Between Steps and Runs

An agent should persist working state for the current task in a structure it writes deliberately, keep durable knowledge in a store with an eviction policy, and treat conversation history as disposable. Anything persisted must be human-readable and provenance-tagged, because a later run trusts it. The test: a run resumes on another machine from the store alone.

Provision a Service Inside an Agent Session

An agent can provision a working Cloudinary environment mid-session by running one unauthenticated npx command, which writes credentials to a local env file and prints a claim URL. The environment expires after 24 hours unless a person claims it, and claiming keeps the same credentials, so nothing written against it needs rewriting.

Reconstructing What an Agent Did After the Fact

To reconstruct an agent run, record one trace per task with a span per model or tool call, log every tool call's arguments alongside its result, store the prompt actually sent rather than the template, and instrument to the OpenTelemetry generative-AI semantic conventions. Replaying the model is a new run; the tool sequence is what replays exactly.

How Much Should an Agent Plan Before Acting?

Plan only far enough to expose assumptions, dependencies, irreversible actions, and review points. Let the agent act directly on cheap, reversible work, preserve an explicit plan for human objection, split only independent subtasks, and re-plan when evidence shows the current route has failed rather than after every successful step.

Agent Permission Models: What Runs Without Asking

A permission model for an agent decides which operations run unattended and which pause for a human. The workable axis is reversibility, not trust: reads and additive writes run freely; deletes, publishes and payments prompt. Enforce it by allowlisting tools per session, scoping and time-bounding credentials, and keeping an escape hatch a human can reach mid-run.

Add circuit breakers to AI agent tool calls

Put a separate circuit breaker around each meaningful tool dependency. Count only availability-related failures, open the breaker when a recent threshold is crossed, reject new calls during recovery, and let only a few probes through half-open. Keep retries bounded behind the breaker, with backoff and jitter, so they cannot recreate the outage.

Compensating actions for multi-step AI agents

Treat every agent side effect as a durable forward-and-undo record: save the result, compensation inputs, and authorization before continuing. Declare the point of no return, make everything after it retryable or manually resolvable, and run compensation as an idempotent, checkpointed workflow with explicit escalation when an undo cannot finish.

Keep an Agent's Context Useful on Long Runs

Keep long agent runs useful by shrinking the always-loaded prompt and tool surface, admitting only relevant results, and moving editable conclusions into durable state. Compact history only at controlled checkpoints, treat prompt caching as a cost measure, and detect context failure with instruction-following evaluations rather than waiting for an overflow error.

Correlate agent runs, model calls, and tools

Use a business operation ID to join retries and replacement runs, and a trace ID to join telemetry within each distributed execution. Propagate both explicitly through model, tool, and queue calls. Keep every identifier opaque and free of customer data so logs and headers remain safe.

Attributing Agent Cost and Latency to Work

Cost attribution for agents means tying every model call back to the unit of work that triggered it. It is a tracing problem: propagate a task identifier into each call, record it as a span with token counts and cost, and account for cached input tokens separately from uncached ones.

Dead-letter queues for failing agent jobs

A dead-letter queue stops repeatedly failing agent jobs from exhausting the normal retry path. Each record must preserve the original payload, the complete failure history, and a correlation ID. Operators should replay only after changing the payload or repairing the failed dependency; otherwise the job is expected to fail again.

Durable Execution for Long-Running Agent Workflows

Durable execution persists the position of a running workflow so a crash resumes from the last completed step rather than restarting the sequence. It is bought with a determinism constraint on workflow code, configured per activity for retries and timeouts, and is worth adopting only when losing a run midway is unacceptable.

Agent Sandbox Egress: Troubleshooting Guide

Filesystem and process isolation alone does not stop an agent from exfiltrating data, because one outbound request is enough. Default-deny egress with a host allowlist, DNS included, enforced outside the sandboxed process, plus short-lived narrowly scoped credentials, fixes more risk per hour of work than a stronger isolation tier.

Expiry semantics for agent-created resources

A resource an agent creates on a user's behalf expires unless a person claims it. Cloudinary's provisioned product environment deletes itself 24 hours after creation, and claiming requires an email address plus a confirmation from that mailbox, so an unattended run leaves lost work rather than orphaned assets and silent billing.

Set a Fallback Policy for AI Agent Model Failures

Use fallback only for retryable transport failures, within the run's deadline. Refusals, invalid tool requests, and contract violations stay on the failure path. A fallback is eligible only when it preserves tool and output contracts and can hold the context required to resume safely.

Golden Sets vs Model Judges in CI Evaluation

A golden set gives a stable, auditable score but costs human effort to build and maintain; a model judge scales to any volume while adding a second non-deterministic system that needs its own validation. Use a small golden set to calibrate the judge, then let the judge cover the rest.

Human Approval Gates for Agents Without Stalling the Run

An approval gate is a pause-and-resume problem before it is a UX one: persist the run so it can wait, batch the questions so a reviewer reads them, show the diff rather than the intent, expire unanswered requests to deny by default with a visible surface, and store the approval record with the changed artefact.

Idempotency Keys for Agent Side Effects

Protect every side-effecting agent action with one idempotency key tied to the intended operation, persist that key across retries, reject changed payloads, and cache the completed result for at least the full retry window. A retry then returns the original outcome instead of repeating the effect.

IP-Locked Delivery Breaks Agent Media Silently

A claimable Cloudinary environment locks media delivery to the public IP that ran the provisioning command. Uploads and transformations still succeed, so the failure surfaces only when a browser on another machine, or a preview deployment, requests the asset. Pass extra addresses at provisioning time, or claim the environment to remove the restriction.

Operational kill switches for autonomous agents

An operational kill switch is an independent control path that stops an autonomous agent from starting or continuing unsafe work. Define measurable thresholds, decision owners, and switch scope before deployment; stop admissions, revoke tool credentials, contain active runs, preserve evidence, maintain a non-agent fallback, and require explicit evidence before redeployment.

Metered billing when an agent drives the workload

Metered vendor billing charges a single unit against several different resources at once. With Cloudinary, one credit buys 1,000 transformations, 1 GB of storage, or 1 GB of delivered bandwidth. Agent-driven work exhausts whichever axis it touches, and transformation and bandwidth totals keep counting for thirty rolling days.

Set Retry Budgets for Agent Tool Calls

Set one retry budget across the entire agent run: cap total attempts, elapsed time, and side-effecting calls. Retry only transient failures on idempotent operations or calls protected by an idempotency key. When any limit is spent, return a typed failure or escalate to a human; never let the model start another loop.

Rotate Secrets Without Breaking Autonomous Agents

Rotate agent secrets by keeping old and new credentials valid during a measured overlap, switching agents to resolve a stable secret reference at each call, and revoking the old credential only after logs show it is no longer used. If dual validity is impossible, drain every run before replacement.

Safety controls for browser-automating agents

Treat every page as untrusted input, enforce origin and download policy outside the model before launch, separate preparation from side effects, and require a human confirmation immediately before purchases, messages, deletions, or credential entry. Validate the browser state around each action and stop when the observed state leaves the approved task.

Sandboxing code an agent wrote

Model-generated code has to run somewhere, and the choice of isolation sets your blast radius. This page compares containers, microVMs, WebAssembly, and hosted sandboxes on startup latency, filesystem and network scope, credential exposure, and operational cost, then states which one each workload profile actually needs.

How to Sandbox Code Run by AI Agents

Run agent-requested code in an isolated environment with no ambient credentials, a minimal filesystem, default-deny network access, and task-scoped capabilities. Enforce separate CPU, memory, wall-time, process-count, output-size, storage, and network limits. Treat the process boundary and syscall filtering as layers, then test every denied path and cleanup action.

Serverless vs long-running hosts for agent workloads

Long-running hosts suit agents better on most axes: they bill for machines rather than wall-clock idle, impose no turn deadline, keep model connections and state warm across turns, and absorb bursts without cold starts. Serverless wins for short, bounded, spiky tool calls where per-account concurrency throttling is acceptable.

Structured Output or Free Text for Agent Results

Use schema-constrained output when software consumes an agent’s result, free text when a person consumes an open-ended answer, and a hybrid when both do. Schemas make malformed results detectable at the boundary; strict rejection preserves that signal, while an unparsed reasoning field retains useful explanation without becoming a hidden dependency.

Sub-agents or One Agent Context?

Keep one agent when the task depends on shared context or produces output the parent cannot verify. Delegate search and summarisation when exploration is large but the return is small, source-backed, and checkable. Limit parallel workers to the provider’s available rate budget.

How to Evaluate Non-Deterministic AI in CI

CI for a non-deterministic system works when you stop asserting on exact output. Assert on properties that survive rephrasing, run each input several times, and gate on a pass rate against a tolerance the product owner sets. Temperature zero narrows variance but never removes it, so design for spread.

Trace Sampling for High-Volume Agent Systems

Use head sampling when predictable overhead matters and rare outcomes are not the selection target. Use tail sampling when slow, failed, or costly traces must survive, accepting complete-trace buffering. In either design, propagate the same sampling decision through every service so retained traces do not arrive with missing downstream spans.

Unauthenticated account creation for AI agents

Cloudinary's agent account-creation endpoint accepts an unauthenticated POST because an agent setting up a new user has no credentials to send. It returns account details, product environment credentials and a machine-readable next-steps block, but those credentials stay inert until the human recipient verifies the account by email and sets a password.

Validate Tool Results Before Model Context

Validate every tool result against a declared success-or-failure schema before it reaches model context. Reject malformed and oversized values as typed failures, never success-like prose. Apply truncation only after the original result passes validation, then shape the validated data to the context budget and record the boundary decision.

Where an agent should put a credential it was issued

A freshly issued credential belongs in a local environment file that version control already ignores, never in the conversation, which is copied into logs and evaluation sets by default, and never in a client-side bundle. Display-time redaction is cosmetic once the value is on disk, so find the vendor's rotation call before you need it.