Desk campaign report · DARPA Lift Challenge

A Paper Helicopter at the DARPA Lift Challenge

In July I spent four days designing a heavy-lift drone that I had no intention of building, using an AI coding agent as the engineering department. In August, ninety real teams flew the same problem in Dayton. This is the recap: what the simulation campaign predicted, what actually happened, and what the comparison says about designing aircraft by conversation.

Author
Cole Ingraham
Date
August 10, 2026
Hardware used
None
Competition
Dayton, OH · Aug 2 to 9

1The challenge

The DARPA Lift Challenge asked a deceptively simple question: how much can a small drone carry? The rules gave it teeth. The aircraft must weigh no more than 25 kg (55 lb) at weigh-in, including fuel or batteries. It must lift a payload of at least 110 lb, take off and land vertically, and haul that payload around a five nautical mile circuit in under thirty minutes. The score is the ratio of payload weight to aircraft weight, judged continuously: 2:1 qualifies, and 4:1 was the headline target DARPA dangled in front of the field, with $6.5 million in prizes behind it.

A 4:1 ratio sounds modest until you sit with the physics. At 4:1 the aircraft itself is twenty percent of the gross weight, so structure, propulsion, energy, and the payload attachment must together weigh less than the margin most aircraft reserve for structure alone. Hover power scales with thrust to the three halves, so every kilogram you fail to shave costs you twice: once on the scale and again in the motor and energy needed to lift it. Production heavy-lift drones, for reference, hover around 1:1.

The finals ran August 2 to 9 at the National Museum of the U.S. Air Force in Dayton, Ohio. Around 480 teams applied, 124 were invited, and roughly ninety flew.

2The experiment: an agent at the drafting table

I was never going to build this aircraft. I have no shop, no thrust stand, and no ambition to certify a hundred-kilogram pendulum swinging under a home-built helicopter. What I had was a question: how far can computational simulation get, when the engineering labor is supplied by an AI agent working under explicit discipline? So I wrote a project charter, handed it to a coding agent (Claude Code), and ran the whole thing as a gated design campaign. The rule was that the agent does the modeling, coding, and analysis; I make the judgment calls at the gates; and nothing advances until the previous gate passes.

4
days, charter to CAD
292
tests in the suite
5
real aircraft calibrated
0
parts purchased

The campaign ran in phases, each with an exit gate:

After the gates closed, the same pipeline produced the artifacts a real program would want: a white paper in which every number is machine-generated from the model (no hand-typed results anywhere), a two-tier bill of materials reconciled against real vendor parts, and a dimensioned CAD ladder from layout skeleton to a printable assembly, with mass properties and gauge checks along the way. The discipline mattered more than any single model: dated findings, written falsification criteria, guard tests against numeric drift, and a standing instruction that the agent brings problems to me for review rather than quietly revising its own conclusions.

3What the desk campaign concluded

The central output was not a design but a ladder. The model scores any configuration as a function of a structural aggressiveness knob, where 1.0 means fielded, production-grade structure and smaller values mean progressively more aggressive competition builds. On that ladder:

A robustly built heavy-lift VTOL qualifies for nothing. Fielded structure tops out at 0.64:1 for piston and 0.99:1 for electric. Reaching 2:1 takes an aggressive competition build (knob around 0.5). Reaching 4:1 requires structure around 0.2: record territory, four to five times lighter per unit gross than any production drone, at the minimum-gauge floor where every panel is as thin as it can physically be made.

At that record-territory setting, the model's two best corners were a piston single-rotor helicopter with cyclic control at 3.93:1 and an electric differential-thrust multirotor at 3.78:1. Note that both sit just under the prize line. The model's honest reading of the challenge was that 4:1 was set slightly beyond reach, which is presumably exactly where a prize line should be.

Getting there required the process to catch its own mistakes, which it did at least once in a way I found genuinely persuasive. The Phase 2 search initially crowned an electric single-rotor helicopter at 4.26:1, the best score of the whole campaign. A Phase 3 audit noticed that this configuration needed a high-torque, low-speed rotor drive that the electric drivetrain model was not paying for, while the piston helicopter and the multirotor both paid their equivalent costs. Pricing the missing transmission dropped the score to somewhere between 2.4 and 3.3, below both rivals. The champion was an accounting error. The correction, not the original search, produced the final recommendation.

The side findings sharpened the picture. Hybrid generators lose to pure electric unless they exceed 3.95 kW/kg, which nothing credible does. A small wing does not pay for itself on a course this short. Jettisoning battery mid-mission is a wash. Thrust-vane control fails quantitatively on descent, arresting nine to twenty-two degrees of payload swing when a plausible gust induces nineteen to forty-five, which is why the recommendation specifies a swashplate. And the bill of materials delivered a warning I want to keep: substituting the real, purchasable version of every part that exists pushes the piston build to 28.5 kg and the electric to 29.9 kg, both over the 25 kg cap. At this scale, the catalog does not contain the aircraft. Whoever builds one must fabricate nearly everything.

Side-view layout schematic of the modeled piston helicopter 3.58 m main rotor 2.17 m boom arm 1.11 m engine 201 mm fwd of mast tail rotor 0.40 m
Figure 1. The modeled prize article at the record-structure corner, every dimension derived from the sizing model rather than drawn by hand: a 3.58 m single main rotor over an open-truss frame, conventional tail rotor, and the engine shifted 201 mm forward of the mast to put the center of gravity under the rotor shaft. Demanded burst power at this corner: 15.4 kW.

Two caveats rode along, clearly flagged. First, several failure modes (the structural weigh-in, sustained battery discharge, vane authority under a real rotor wake) can only be measured on a bench I do not have, so they were carried as unretired risks rather than retired ones. Second, a late analysis found that the model's assumed cruise efficiency was badly generous for a bare rotor at the course's low speeds, a correction that would push scores meaningfully down. I chose to flag it for reconciliation rather than apply it silently. Both caveats matter below.

4Then reality flew

The competition concluded on August 9. The result that matters most is a single fact: nobody scored 4:1. The prize threshold went unclaimed by a field of ninety teams that included funded startups and university groups with actual wind tunnels.

#TeamAircraftPayloadRatioDesign, as reported
1AVIDrone (Columbia, MD)13.2 kg50.8 kg3.84:1electric tandem-rotor
2MTech Operations (Sudbury, MA)14.6 kg53.1 kg3.64:1electric, configuration unreported
3Xtreme Aerial Concepts (San Jose, CA)24.9 kg85.9 kg3.45:1helicopter, custom turboshaft
4H-Squared (Boston, MA)24.9 kg73.5 kg2.96:1unreported
5MacGyver (Baltimore, MD)24.9 kg62.1 kg2.51:1unreported
Final top five. Figures reconstructed from broadcast telemetry by an unofficial community mirror and cross-checked against observer posts; treat them as close but not official. One early entry posted 9.63:1 and was delisted within two days, which tells you something about both the rules and the physics.

The podium splits into two strategies. The winners went small: AVIDrone and MTech both flew aircraft around thirteen to fifteen kilograms, barely half the allowance, and both were electric. Third place went the other way: Xtreme Aerial Concepts arrived at 24.9 of a permitted 25 kg with a single-main-rotor helicopter built around a custom turboshaft from Jakadofsky, a small turbine rated near 22 horsepower that weighs about 3.6 kg. That machine carried 85.9 kg around the course, the heaviest payload flown all week, and reportedly attracted a $20 million seed round before the awards ceremony.

5The scorecard

The interesting comparison is not whether the model named the winner. It is whether the model's map of the possible matched the territory. Here is the whole story on one axis:

Dot plot comparing modeled scores with flown competition scores on a shared payload ratio axis 2:1 qualifies 4:1 prize 0 1 2 3 4 payload : aircraft weight MODELED FLOWN fielded structure aggressive (0.5) record (0.2) 3.45 3.84 3.93 fuel-burning electric n/a
Figure 2. Modeled score corners (top row) against the flown top five (bottom row). The whole podium landed between the model's two record-structure corners and the qualifying line, and nothing crossed 4:1. Hover any point for details; the full numbers are in the tables.

What the competition validated. The headline held: 4:1 went unclaimed, and the best flown score, 3.84, sits neatly between the model's two record-structure corners of 3.78 and 3.93. The ladder itself calibrates well. Inverting the model (asking what structural knob setting reproduces each real score) places the podium at 0.19 to 0.24 and fourth place at 0.32: the best teams on earth showed up with precisely the record-territory structure the model said the prize demands, and it was still not quite enough. The architecture call held too, in class if not in rank: the heaviest lifter was a fuel-burning single-main-rotor helicopter riding the weigh-in cap, which is the strategy and layout the model recommended for the prize corner.

The third-place machine deserves its own paragraph, because it is the closest thing to the model's design flying in the world. The model demanded 15.4 kW of burst shaft power at the prize corner; Xtreme flew a 16.4 kW engine, a match within six percent. The model's gross weight at the inverted knob setting comes out at 111.3 kg against their real 110.8. A pre-event analysis of that engine-airframe pairing estimated a rotor between 3.7 and 4.6 m; the model drew 3.58. Down the row of the specification, a simulation that never touched hardware and a company that just raised $20 million landed within about ten percent of each other.

What it invalidated, or at least embarrassed. Three things, in increasing order of interest:

  1. The engine axis was too narrow. The charter enumerated electric, piston, and hybrid. Nobody wrote down "turboshaft," so the agent never considered one. The irony is that the campaign had already derived the reason turbines win here: it proved that course energy is nearly free on a thirty-minute mission, which means specific power is everything, and a small turbine delivers roughly 4.5 kW/kg against a piston engine's 1.3. The logic was sitting in the findings, unapplied to the axis it should have widened. Third place applied it.
  2. Fuel versus electric at the top came out backwards. The model ranked the piston helicopter first by a small margin (3.93 over 3.78); reality put two electrics above the turbine helicopter. The margins are inside the model's own stated sensitivities, but the sign was wrong.
  3. The winners flew half the cap, and the model could not have suggested it. The formulation fixed the aircraft at 25 kg and maximized payload, and its calibrated structure model scales linearly with gross weight, so building smaller offered no modeled advantage. AVIDrone won at 13.2 kg. The likely mechanism is one the parts list had already found from the other direction: at 25 kg scale the catalog holds no adequate parts, but at 13 kg the commercial ecosystem is rich. The optimum was set by the supply chain, a variable the physics never sees.

There is also a quieter reconciliation datapoint. That flagged cruise-efficiency correction, had I applied it, would have dragged the predicted corners well below where the field actually landed. The uncorrected model matched reality; the correction would have spoiled it. That is worth understanding rather than celebrating (real teams likely flew faster than the modeled speed, and battery mass at these power levels is set by power draw rather than energy anyway), but it argues the flag deserved its provisional status.

QuantityModeledFlown (3rd place)Delta
Architecturesingle-rotor heli, fuelsingle-rotor heli, fuelmatch
Weigh-in strategyride the 25 kg cap24.9 kgmatch
Burst shaft power15.4 kW16.4 kW rated6%
Gross weight111.3 kg *110.8 kg<1%
Main rotor diameter3.58 m3.7 to 4.6 m (est.)in band
Engine type2-stroke pistonturboshaftmiss
Score3.93:13.45:112%
Model versus the third-place aircraft. The starred gross weight is evaluated at the knob setting that reproduces their score, so its agreement is partly mechanical; the power, rotor, and strategy comparisons are not.

6Where agentic design goes from here

I want to be careful about the claim here. A desk campaign did not out-design ninety teams; it did not design a flyable aircraft at all, and the bench-only risks it flagged remain exactly as unmeasured as the day it flagged them. What it did do, in four days and for the price of some tokens, was draw a map of the feasible region that the subsequent competition traced almost exactly: the ceiling, the architecture of the heaviest lifter, its power and geometry to within roughly ten percent, and the structural regime required to stand on the podium.

The errors are at least as instructive as the hits, because none of them were arithmetic. Every miss was a question the campaign never asked: an engine type absent from a human-written list, a supply-chain effect outside the physics, a formulation that pinned a variable reality treated as free. The agent executed the charter faithfully, corrected its own accounting when audits caught it, and never challenged the charter's boundaries. If I ran this again, the discipline I would add is a standing adversarial pass over the problem statement itself: enumerate the axes you were given, then spend real effort arguing that the list is wrong. That is a prompt away, which is rather the point.

So here is what I think this style of work is actually for, near term:

The unclaimed prize is the right closing note. The model said 4:1 sits just past the edge of record-territory structure, and ninety teams spent a week in Ohio confirming it. Somewhere in next year's field, someone will pair a turbine with a lighter truss and take the model's remaining margin away. When they do, I hope they checked their assumptions against a simulation first, and I hope the simulation had the good manners to argue back.