We use cookies for analytics and to measure our Google Ads. Nothing is loaded until you accept. Cookie policy

Agent Runs · Swarm Lab

Work that sounds impossible. Done by a swarm.

A flagship Agent Run splits one enterprise-scale job into thousands of small agent tasks: reading every contract, screening every supplier, mapping a whole sector. Supervisor agents re-check the work, a person signs off every risk decision, and you get the result with an evidence log. Prices are fixed before the run starts.

Last updated:

Impossible → delivered

Seven flagship runs we scope today. Flip a card to see how the swarm does it.

  • F1

    Read every contract you’ve ever signed — by Friday.

    How the swarm does it

    Worker agents read each contract and pull out parties, dates, renewal terms, liability caps and change-of-control clauses. Supervisors re-read a sample of every worker’s output. An editor builds the risk register.

    Scale
    ~20,000 contracts
    Run time
    72 hours
    Quality check
    A double-read sample; disagreements go to your legal team. The error rate found is stated in the report.
    From €40,000
  • F2

    Find every invoice you paid twice in the last three years — and up to ten where records exist.

    How the swarm does it

    Agents normalise your payment history across all entities (three years as standard, up to ten where the records exist), match on amount, supplier, date and reference variations, and assemble each suspected duplicate with both payment records.

    Scale
    3 years (up to 10), all entities
    Run time
    1 week
    Quality check
    Every candidate comes with its evidence. Your AP team confirms each one before any recovery claim.
    From €30,000or on contingency
  • F3

    Document a system nobody who built it still works on.

    How the swarm does it

    Agents read the VB6, COBOL or Access code module by module, write what each part does and which data it touches, and draft a migration plan from the dependency map.

    Scale
    The whole code base
    Run time
    2 weeks
    Quality check
    A second agent checks each module’s documentation against the code; your developers review a sample and the plan.
    From €60,000
  • F4

    Screen your whole supplier base in 48 hours.

    How the swarm does it

    Agents check every supplier against sanctions lists, VIES, insolvency registers and the press, and write a short finding per hit with its source.

    Scale
    ~20,000 suppliers
    Run time
    48 hours
    Quality check
    Every hit links to its source. A person decides whether each match is real.
    From €25,000
  • F5

    Map every company in your sector across the EU — in a week.

    How the swarm does it

    Agents collect public data on each company, score it against criteria you agree upfront, and cite where every data point came from.

    Scale
    ~50,000 companies
    Run time
    1 week
    Quality check
    A sample is checked by a person, and the scoring rules are written down before the run starts.
    From €20,000
  • F6

    Turn every EU and Luxembourg rule that applies to you into one register.

    How the swarm does it

    Agents read the texts that apply to your activities and list each obligation with the article it comes from, who owns it and how often it applies.

    Scale
    Every applicable EU and Luxembourg text
    Run time
    10 days
    Quality check
    Every obligation cites its article. A lawyer — your counsel or a firm you choose — signs off before delivery.
    From €30,000
  • F8

    Sweep 50,000 patents and papers in five days.

    How the swarm does it

    Agents read each document, keep the passages relevant to your claims and rank them, so your attorney reads dozens, not thousands.

    Scale
    ~50,000 patents and papers
    Run time
    5 days
    Quality check
    Each result shows the exact passage that matters. Your patent attorney makes the call.
    From €25,000

These are offers, sized from our demo runs. No flagship run has been delivered yet; when one is, its measured accuracy is published below.

Proof from the frontier

What large agent campaigns have already done in public, with the numbers as the labs reported them.

Public results from AI labs — not 20 More projects.

  1. , Anthropic

    A C compiler written by 16 parallel agents

    A team of 16 Claude agents wrote a 100,000-line C compiler, in Rust, that can build Linux 6.9 on x86, ARM and RISC-V.

    16
    agents in parallel
    ~2,000
    Claude Code sessions
    ~$20,000
    API cost

    Caveat: A research project by Anthropic’s own engineers, not a client delivery.

    Source: Anthropic Engineering — Building a C compiler with parallel Claudes

  2. , Anthropic

    Project Glasswing: vulnerabilities in 1,000+ open-source projects

    Claude Mythos Preview scanned more than 1,000 open-source projects. Independent security firms assessed 1,752 of the findings it rated high or critical.

    23,019
    findings in total
    1,752
    high/critical findings assessed
    90.6%
    of those were real

    Caveat: The 90.6% applies to the 1,752 independently assessed high- and critical-rated findings, not to all 23,019.

    Source: Anthropic — Project Glasswing: initial update

  3. , Anthropic

    A new enzyme system found by a ~950-session campaign

    Claude agents searched protein data for about 21 hours and found a previously uncharacterised enzyme system with CRISPR-like repeats, called ART.

    ~950
    agent sessions
    ~21 h
    of searching
    210M
    tokens

    Tap a step to see what happened there.

    Caveat: Human scientists did all the lab work, and the new system’s function is not yet known.

    Source: Anthropic — Claude discovers a novel enzyme system

Our campaigns

Our first flagship case studies publish here with measured accuracy.

Questions about flagship runs

Does a flagship run mean thousands of agents at the same time?

No. It means thousands of agent tasks. At most about 50 agents work at once, in batches, until every task is done and checked.

Have you delivered a flagship run yet?

Not yet. The runs on this page are offers, sized from our demo runs. The first measured results will be published on this page.

How do you measure accuracy?

A person checks a random sample of the output against the source documents. The report states the error rate found in that sample and how the sample was drawn.

What does a flagship run cost?

Flagship runs start at €20,000, with a fixed price agreed after a scoping call. Each card shows its own starting price.

Where does our data go?

Models run in EU regions (Claude through Amazon Bedrock). The run can execute inside your own cloud account, and we never store your passwords. We sign a GDPR data processing agreement before anything starts.

Scope a flagship run

Four questions. We reply within two working days with what a run would do, how long it would take and a fixed price.

From €20,000