← All work

The Discovery Pod Operating System

A creator sourcing team was assigning work using speed estimates that were wrong for every single person, in both directions, by up to 2.4x. Nobody knew, because nothing measured it.

Client
UK influencer marketing agency, health and wellness
Engagement
Full-stack platform: database, application, AI assignment engine, migration
Scale
14 clients, 8-person pod, 20+ concurrent campaigns
Findings measured
First 6 weeks of operation, June to August 2026

Four things the agency could not previously have known.

Each had been silently costing money for as long as the function existed.

Finding 01

Every speed estimate was wrong, in both directions

Work was assigned against speeds the team believed to be true. Not one was right, and critically they were not uniformly optimistic or uniformly pessimistic. They were simply unrelated to reality.

Assumed pace against measured pace, per available working hour. Measured across 31,294 finds and 65,506 vetting decisions.
Team member Discovery Error Vetting Error
A56 → 1342.4x under100 → 1821.8x under
B50 → 901.8x under150 → 1701.1x under
C50 → 701.4x under150 → 1270.85x over
D60 → 681.1x under100 → 2092.1x under
E40 → 441.1x under120 → 940.79x over
F60 → 410.68x over100 → 98roughly right
Gvetting only150 → 960.64x over

Pod-wide, discovery capacity was under-estimated by about 40% while vetting capacity was over-estimated by about 18%.

When estimates are wrong in both directions, work is not merely mis-sized. It goes to the wrong people.

Someone believed to be a fast vetter and loaded accordingly was in fact one of the slower ones, while a genuinely fast discoverer sat under-used. Every capacity decision made before this point, whether a campaign could be accepted, whether a deadline was reachable, whether another hire was needed, rested on it.

Finding 02

The survival assumption was under-provisioning every campaign

Planning assumed 75% of discovered creators would survive quality vetting. The measured rate across 68,942 real verdicts is 64.8%. That gap is not academic. It has an exact operational cost.

To deliver 1,000 surviving creators, the plan sizes discovery at 1,333 finds. At the real survival rate those yield 864. Hitting 1,000 actually requires 1,543 finds, 210 more than the plan provisions, every time.

Across the six-week window the wrong assumption accounts for roughly 9,400 vetting decisions the plan never budgeted for. That shortfall was previously absorbed as unplanned top-up work and missed targets, which read as effort problems rather than arithmetic problems.

Finding 03

A 28-point quality spread between researchers, previously invisible

Two people producing at similar volume can differ by 28 points on whether that volume is worth anything. The difference lands entirely on the vetters downstream, as hours spent rejecting work that should never have been submitted.

Survival rate by researcher, across 29,167 peer-vetted finds.
ResearcherFinds vettedSurvival rate
A65291.6%
B3,07786.7%
C2,76578.5%
D10,96476.3%
E8,03563.8%
F3,67463.3%

Before the platform, every one of them looked identical on a tracker. The same measurement runs per campaign type, where survival ranges from 33.9% to 100% across 28 categories, which is the direct evidence that a single flat assumption cannot be right for every brief.

Finding 04

About a third of all rejections were objectively preventable

Rejections are captured as structured reasons rather than free text, which makes them countable. Of 23,777 tagged rejections, 8,380, roughly 35%, failed on objective criteria: follower count, region, age range, language, inactive accounts, business pages.

All of those are knowable before anyone opens the profile. They represent work that was sourced, submitted, reviewed and thrown away.

8,380

Objectively preventable rejections in six weeks

310

Paid hours consumed sourcing and rejecting them

~20%

Of total pod capacity, spent on avoidable rework

28p

Cost to deliver one vetted creator, now calculable

The function could be reacted to, not planned.

The old process ran on spreadsheets and one person's judgment. Assignment was gut-feel. Progress was a number someone typed into a cell, which drifted from reality immediately and was then used to plan the following week.

Quality was entirely invisible. Peer review happened, but nothing measured whether it caught anything, or whose work kept failing it. State lived per campaign, so the same creator was rediscovered repeatedly, and a creator rejected for a client could resurface for that same client a month later.

Every action is a timestamped event. Throughput, quality and trust are derived live from the work itself.

Nobody maintains the planning numbers, because the planning numbers are a query over what actually happened. That is the entire reason the four findings above exist. They are not a reporting feature bolted on. They are a by-product of people doing their normal work on a system that records it properly.

One platform, three surfaces.

The assignment engine, in two layers, in this order

A deterministic layer computes what is physically possible: spare capacity per person, work spread to the deadline, the peer-review exclusion chain, a quality floor on who may do final vetting. A language model then judges the shortlist on fit and context. Then the deterministic layer clamps the model's answer back to capacity. That order is the whole design. The failure mode of an AI assignment tool is confident overcommitment, and the fix is not a better prompt. It is refusing to let a model be the final authority on arithmetic.

The unit of state is the pairing, not the creator

A creator is global identity, their state is per client. This makes cross-researcher deduplication structural rather than policed, lets the same creator be live for one brand and rejected for another, and makes a rejection permanent for that client.

Progress is a query, never a counter

Completion, quality scores, survival rates, latency and throughput are all derived from events on read. They cannot be gamed by editing a cell and they cannot drift.

The system says when it does not know

After an incident where a failed query rendered a confident "queue clear," every count carries a trust state. A number that cannot be verified reads "couldn't load, retry." A number that cannot be reconciled says so rather than quietly showing a wrong total. An operations tool that is confidently wrong is worse than one that is honestly unavailable, because the entire value of the tool is that people stop double-checking it.

The number every campaign can now be priced against.

Cost to deliver one vetted creator: roughly 28p to 42p. Sourcing, vetting, and absorbing the third who do not survive. Before this platform the agency could not have calculated that figure at all, because neither the pace nor the survival rate was measured.

Avoidable rework: £2,500 to £3,700 in the six weeks measured, on the order of £21,000 to £32,000 a year if the pattern holds. Treat that as the ceiling on what tighter briefs and source-side filtering could recover rather than a guaranteed saving. But the ceiling is the size of the prize, and it was invisible before.

Capacity allocated to the right people. The platform did not make anyone work faster. It revealed that discovery capacity was under-estimated by 40% while vetting capacity was over-estimated by 18%, which means the pod had been systematically over-committing one stage and under-using the other. Correcting that is free throughput.

What this case study deliberately does not claim.

The measurement exercise behind this page was designed to disprove claims as readily as support them. A case study that claims everything is a case study nobody should trust.

That the platform made the team faster

There is no measured record of true pre-platform throughput. The old system stored planning assumptions and one real work record. What can be proven is that guesswork was replaced with measurement, and the guesswork was wrong in every case. That is a planning-accuracy result, not a productivity claim.

A duplicate-work saving

Deduplication runs, but the counter that would record how many duplicate attempts it blocks was never populated. The saving is real and currently unmeasurable, so it is not estimated here.

A client conversion figure

The loop that would record which handed-over creators were ultimately taken on was built but never fed. Five campaign opt-in rates exist, which is too small a sample to price anything against.

That the cost figures are precise

They rest on two inputs from the agency: a loaded hourly cost given as a range, and an estimate of how much of a paid hour is hands-on work. Both are stated as ranges for that reason. The underlying pace and survival figures are measured regardless of what is assumed about either.

Most agencies cannot say what one vetted creator costs.

Because nothing measures it. Book a call and we will work out your number, and how many hours a week your pod is losing to work that never had to happen.

Book a Call