Choosing a Coding Assistant for a 20-Person Team
For a twenty-engineer team, buy per-seat only when usage is even across the team; buy usage-metered when a few engineers drive most calls, which is more common. Screen SSO, audit logs, residency and retention first, run a fixed-task trial that measures completion and rework, and set a review date.
Why twenty engineers is the awkward size
A five-person team picks a coding assistant the way it picks a terminal emulator: whoever cares most installs it, the rest follow, and the cost is a rounding error. A two-hundred-person organisation runs procurement, and procurement, for all its friction, forces the security questionnaire, the legal review and the budget line to happen before anyone commits.
Twenty engineers sits between those. It is past the size where one person’s preference can decide — there are now at least three editors in use, two opinions about agents versus assistants, and a manager who will be asked to justify the invoice. It is below the size where a formal procurement process pays for itself: nobody is going to spend six weeks and a lawyer on a tool that costs less than one recruiter fee. So the decision gets made the way it gets made at five people, with the consequences of making it at two hundred. That is the mechanism behind most bad choices in this space, and it is worth naming before comparing anything, because the fix is not to pick a better tool. It is to borrow the four or five procurement questions that actually matter and answer them in an afternoon.
The rest of this page weighs the realistic options against those questions. It deliberately does not rank vendors: capability and pricing both move on a cadence of months, and any table of names here would be stale before it was indexed.
The options
At twenty engineers there are three shapes on the table, and the third is the one teams forget to write down.
Option A — a seat-licensed assistant. One price per engineer per month, usually with a soft or hard cap on model calls, an admin console, and an IDE extension as the primary surface. The pricing assumes each seat consumes roughly what a typical user consumes.
Option B — a usage-metered assistant or agent. Billing follows tokens, requests or task minutes. Often terminal-first or agent-first, sometimes with an IDE surface layered on. The pricing bills the actual distribution of use across the team, whatever shape it has.
Option C — leave it as it is. Several engineers already have something installed on personal or expensed plans; the team standardises nothing. This is what most twenty-person teams are running today, and it has real costs that only show up when the first security question arrives.
Each criterion below states each option’s standing. The criteria are the ones that decide the outcome for a team of this size, and only those.
Criterion 1: seat pricing and usage pricing invert at a knowable point
Seat and usage pricing are not two prices for the same thing; they are two bets on the shape of consumption.
A seat price is set to cover a typical-to-heavy user, so it is only a good deal when usage is roughly even across the team — everyone drawing on the assistant most days, at similar volume. Under that shape, twenty seats cost less than metering twenty people individually, because the vendor has priced in volume that the team actually uses.
A usage price bills consumption directly. It wins when a minority of engineers drive most of the calls — three or four people running agents through long tasks all day while the other sixteen accept the occasional completion. Under that shape, seats charge you sixteen times for capacity nobody draws on, and metering charges you for the four people who use it.
The second shape is the more common one. Adoption inside a team is not uniform: it follows the engineers whose work suits the tool and whose habits formed early. The honest planning assumption for a team that has not measured itself is a skewed distribution, which points at usage pricing until the data says otherwise.
The inversion point is knowable, not mystical. Take one month of the team’s actual call volume — most vendors expose per-user usage in the admin console, and Option C teams can ask people to export theirs. Sort by user. If the top quarter of users account for well over half of the volume, metering is likely to be cheaper; if the curve is flat, seats are. Illustrative arithmetic, not a vendor quote: if a seat costs S per month, and metered spend for a heavy user works out near 3S while a light user’s works out near 0.2S, then a team of four heavy and sixteen light users spends 12S + 3.2S = 15.2S metered against 20S on seats. Flip the ratio to twelve heavy and eight light and metering costs 37.6S against the same 20S. The break-even is a function of the distribution, and the distribution is measurable in a month.
| Even usage across the team | A minority drive most calls (more common) | |
|---|---|---|
| Seat-licensed (A) | Cheaper | Pays for idle seats |
| Usage-metered (B) | Pays for volume | Cheaper |
| Unstandardised (C) | Unknown — spend is spread across expense claims | Unknown, and usually higher than either, because heavy users are already on top-tier personal plans |
Option C’s standing on this criterion is that it has no standing: the team does not know what it spends. That is not neutral. It is the shape most likely to be paying the seat price for the heavy users and nothing for the light ones, which is the worst combination of both models.
A related question — which model tier the heavy users are actually calling — moves the metered number a lot. Frontier versus mid-tier model selection per task class covers when the expensive tier earns its cost and when it does not.
Criterion 2: editor lock-in is the cost nobody prices
Every comparison spreadsheet has a row for price and a row for capability. Almost none has a row for what it costs to switch, and for a twenty-person team that row is usually the biggest number on the sheet.
The mechanism is muscle memory. After a few weeks with any assistant, an engineer stops thinking about how to invoke it: the keybinding for accept, the gesture for reject, the way a multi-file change is reviewed and applied, where the diff appears, how a session is resumed. Those habits are per-tool, not per-editor. Move the team to a different assistant — even a strictly better one — and each engineer spends days back at conscious-competence speed, and the team as a whole spends weeks. This happens regardless of which tool is better, and it happens twice if the first choice is reversed.
The standing of each option:
- Option A tends to bind hardest, because its whole value is a tight IDE integration. The better the extension, the more habit forms around it. Migrating away means retraining hands, not just reissuing licences.
- Option B, when it is terminal- or agent-first, binds less to a specific editor but still binds to a workflow: how tasks are phrased, how permissions are granted, how output is reviewed. Switching agents costs less in keystrokes and more in relearning what the tool will and will not do unsupervised.
- Option C has already paid this cost, in fragments, and will pay it again the day it standardises — because standardising means at least some engineers give up the tool they built habits on. The longer C persists, the more expensive the eventual move.
The practical consequence is that the first standardised choice is stickier than it looks. That argues for two things: getting the eliminating criteria (below) right before the trial rather than after, so the team does not have to switch off a tool it has already habituated to; and choosing, where two candidates are close, the one whose surface the team already partially uses. It also argues against a long trial of three tools in parallel, which builds three sets of habits and then discards two.
Where the vendor exposes the same operations through both an editor extension and an MCP server, the lock-in calculus changes, because the operations can outlive the editor choice; an IDE extension against an MCP server for the same vendor operations works through that trade.
Criterion 3: enterprise controls belong at the start, not the end
SSO, audit logs, data residency and zero-retention agreements are the requirements most likely to eliminate a candidate late. They are the questions the security lead or the largest customer asks after the trial is finished and the habits are formed, and if the answer is no, the team either switches — paying criterion 2 again — or carries an unapproved tool.
The mechanism is simply that these features are gated by plan tier, and the tiers that carry them are not the tiers a twenty-person team trials on. A team evaluates on the plan it can sign up for with a card, finds the tool excellent, and then discovers that SSO or audit logging or a regional deployment lives on a tier priced for a hundred seats, or requires a sales conversation, or is not offered at all.
So invert the order. Before any engineer installs anything for evaluation, write down which of the four the team actually needs — not might want, needs, because a customer contract or a compliance regime says so — and check each candidate’s tier sheet against that list. Candidates that fail are out before the trial, which is cheap; candidates that fail after the trial are out at the cost of weeks.
Standing by option:
- Option A vendors typically carry these controls, but on their upper tiers. The seat price the team modelled in criterion 1 may not be the price of the tier that passes criterion 3. Re-run the arithmetic on the tier that actually qualifies.
- Option B varies more. Some usage-metered products expose enterprise controls at any spend; some expose none below a committed contract. This is checkable in an hour and must be checked, because metered products are the ones a team is most likely to adopt bottom-up without anyone reading the tier sheet.
- Option C fails this criterion by construction. Personal plans do not offer SSO, do not produce a team audit log, and make no residency or retention commitment to the company, because the company is not the customer.
One caution in the other direction: do not require controls the team cannot articulate a reason for. Requiring all four eliminates most of the market and pushes the price into a range that a twenty-person team cannot defend. Require what a named customer, regulator or policy actually demands, and write that reason next to the requirement.
Criterion 4: retention, training and residency are three commitments, not one
Vendors describe their data handling in a single reassuring paragraph. That paragraph is doing three separate jobs, and the team needs to know which of the three was actually agreed.
- Retention is how long the vendor keeps prompts, code context and outputs after processing. Zero-retention means not stored beyond the request; a thirty-day window means stored and deletable; default retention means stored per the vendor’s general policy.
- Training is whether that data is used to improve the vendor’s or a model provider’s models. A tool can retain data and not train on it, or train on it and claim not to retain it, and both configurations exist in the market.
- Residency is where processing and any storage happen geographically, which is what a customer contract or a data-protection regime actually constrains.
A statement like the vendor does not train on your code answers the second question and says nothing about the first or third. It is common for a team to believe it has a zero-retention agreement when what it has is a no-training clause and a thirty-day log.
Only a dated citation from the vendor’s own terms settles any of the three. Not a sales deck, not a blog post, not a support reply — the terms or data-processing addendum, with the date it was retrieved, saved alongside the decision. Terms change; a decision file that says zero retention with no date and no link is an assertion, not a record.
Standing by option: A and B are on equal footing here, because this is a documentation question, not a product-shape question — either can be checked in the same way. Option C, again, has no commitments to the company at all, because the agreements are between the vendor and individual engineers on personal terms.
Criterion 5: a trial that measures acceptance rate measures usage, not outcome
Most vendor dashboards lead with acceptance rate — the share of suggestions an engineer kept. It is an easy number and it is the wrong one. Acceptance rate tells the team how much the tool is used, which is a fact about habit and about how aggressively the tool suggests. It does not tell the team whether the work got done faster, or whether it needed redoing.
A trial worth running fixes a task set and compares completion and rework across the same work. Concretely: choose a handful of tasks representative of what the team actually ships — a bug with a reproduction, a small feature behind a flag, a refactor with tests, a dependency upgrade. Have engineers do them with each candidate, and with no assistant as the baseline. Measure completion — did the task reach a mergeable state, and in how long — and rework — how much of the assisted output was reverted, rewritten or fixed in review within the following week. Rework is the number vendors never show, and it is the one that separates a tool that produces plausible code from one that produces correct code.
This is harder than reading a dashboard, and it is worth doing precisely because the published evidence does not settle the question. The controlled experiment in Peng and colleagues’ study of GitHub Copilot on a fixed programming task reported the assisted group finishing substantially faster on that one task — a result about completion time under lab conditions, not about rework or about a team’s real backlog. Studies since have pointed in different directions, and this site does not treat the productivity effect as established; the responsible reading is that the effect is task-dependent and must be measured on the team’s own work.
Two framings help keep the trial honest. Martin Fowler’s site carries an argument that developer productivity should be measured through the people doing the work rather than through activity counts, which is exactly the case against acceptance rate as a headline. And the DORA four keys — deployment frequency, lead time for changes, change failure rate and time to restore — are the right place to look for whether an assistant changed anything at the team level over months. Lead time and change failure rate are, in effect, completion and rework measured on the whole pipeline instead of a task set. They will not resolve in a two-week trial, which is why the trial needs its own fixed-task measurement, but they are what the review date (below) should re-examine.
Standing by option: A and B are trialled the same way, and the trial is the point where the price-shape data for criterion 1 also arrives, because the trial produces the first honest per-user usage numbers. Option C cannot be trialled, because there is no controlled comparison against a scattering of personal tools — which is another way of saying it can never be justified, only continued.
How to run that measurement over the longer term, and what to do with the DORA signals once they move, is the subject of measuring the impact of a coding assistant on an engineering team.
Criterion 6: the decision needs a review date
Pricing and capability both move on a cadence of months. Tiers are renamed, caps change, a model generation ships and shifts which tasks the mid tier can handle, a vendor adds the audit log that eliminated it last time. A choice made carefully in one quarter is, two quarters later, an assumption nobody has re-examined.
Write the review date into the decision. Six months is a defensible default: long enough that the lock-in cost of criterion 2 is not paid pointlessly, short enough that a tier change or a competitor’s move gets noticed. At the review, re-check the four commitments of criterion 4 against the vendor’s current terms, re-run the per-user usage sort of criterion 1, and look at whether the DORA numbers moved. If nothing has changed, the review costs an hour and the decision stands with a new date.
Standing by option: A and B are equal, since the review is a habit of the team, not a feature of the tool. Option C has no decision to review, which is the quiet way it persists.
Which to pick when
- Pick a seat-licensed assistant (A) if the team’s measured usage is even — most engineers using it most days at similar volume — and the tier that carries the enterprise controls the team actually needs is priced within reach at twenty seats. Take the tier sheet, not the entry price, into the arithmetic.
- Pick a usage-metered assistant or agent (B) if usage is skewed — a minority driving most calls, which is the more common shape and the right default when the team has not measured — and the metered product either exposes the needed controls at any spend or the team genuinely needs none of the four. Watch the heavy users’ model-tier choices, because that is where a metered bill grows.
- Do not stay with Option C beyond the month it takes to gather usage data from it. It fails the controls criterion by construction, makes no commitments to the company on retention, training or residency, cannot be trialled, and grows more expensive to leave every week it continues.
Whichever of A or B wins, do it in this order: list the controls the team can name a reason for, eliminate on the tier sheet, save dated citations for the three data commitments, run a fixed-task trial that records completion and rework, model the price on the qualifying tier against the trial’s usage distribution, and write the review date into the decision. That is a procurement process scaled to an afternoon, which is the process a twenty-person team can afford and the one it usually skips.
Sources
- Peng and colleagues' study of GitHub Copilot on a fixed programming taskarxiv.org
- developer productivity should be measured through the people doing the work rather than through activity countsmartinfowler.com
- DORA four keys — deployment frequency, lead time for changes, change failure rate and time to restoredora.dev
See also
Stop failing tools from consuming agent time and worsening outages by adding scoped breakers, bounded retries, recovery intervals, and limited probes.
How to record reversible agent actions, mark the point of no return, and make compensation idempotent, resumable, and operable.
A context policy for long agent runs: trim tools, admit relevant evidence, persist editable conclusions, compact carefully, and test drift.
Carry trace and business operation IDs through agent runs, model calls, tools, and queues without leaking customer data.