In July 2026 a test pilot took off from Eglin Air Force Base in an F-16 and flipped a switch that handed the jet to software. DARPA announced on July 16 that an artificial intelligence agent had controlled the aircraft in flight, the first such sortie for the Viper Experimentation and Next-gen Operations Model, or VENOM. The pilot, according to Stars and Stripes, flew the takeoff and then monitored the cockpit while the agent flew.
That arrangement, a fighter flown by code with a human ready to take it back, is now the Air Force’s main method for getting combat autonomy out of the laboratory. It deserves to be understood precisely, because the results are routinely read as proof that software can fight. A surrogate jet with a safety pilot proves something narrower: that a given agent can be flown and evaluated safely on a real airframe. Whether that agent can sense, decide and survive alone is a separate question, and each stage of testing should be made to answer it with published evidence.
What a safety pilot buys, and what it hides
The pipeline began in simulation. In DARPA’s AlphaDogfight Trials of August 2020, eight teams flew simulated F-16s, and the winning agent from Heron Systems beat an Air Force pilot 5-0. Simulation is cheap and repeatable. It also flatters the software. The agent sees a clean model of the world, and nothing it does can break an airplane.
The X-62A VISTA, a specially modified F-16 at the Air Force Test Pilot School at Edwards, exists to close that gap. Under DARPA’s Air Combat Evolution program the aircraft flew 21 test flights between December 2022 and September 2023, ending with two weeks of engagements against a crewed F-16 at speeds up to 1,200 miles per hour and separations inside 2,000 feet. DARPA described the flights in April 2024 as the first in-air tests of AI flying an F-16 against a human-piloted one within visual range. Two pilots rode in the X-62A and, officials said, never had to take over.
The officials running the test were careful about its meaning. Lt. Col. Ryan Hefron, then the program manager, told reporters the purpose was “to demonstrate we can safely test these AI agents” in that environment, and the Air Force declined to say how often the machine won. The demonstrated result was a test method. Tactical superiority was neither claimed nor shown.
The limits were also physical. The War Zone reported in September 2024 that during the dogfights the X-62A exchanged position and other state information directly with its opponent, mainly for safety, and that the jet had no sensor suite of its own able to track an adversary all around it. An agent fed its opponent’s location over a data link has solved the maneuvering problem. Finding and tracking the target, the harder half of air combat, was done for it.
Sensors first, then more aircraft
The work since 2024 has gone after those gaps in order. The X-62A is partway through a Mission Systems Upgrade, and in a campaign called HAVE HEAT it flew 27 autonomous intercepts over eight flights against a T-38, with the agent working from a Lockheed Martin Legion Pod infrared search-and-track sensor instead of shared data, according to an August 2026 account in Military & Aerospace Electronics. A Raytheon PhantomStrike radar is planned. That is a real step: the software had to act on live, imperfect sensor data.
VENOM addresses a different shortage. One research jet at Edwards cannot test several autonomous aircraft together. The Eglin program is modifying six F-16s with a kit that, DARPA says, automates flight controls and sensors without changing the jet’s core software. Those aircraft feed DARPA’s Artificial Intelligence Reinforcements program, which targets multi-ship combat beyond visual range and plans to develop agents on crewed F-16 testbeds before transferring them to an uncrewed aircraft. So far VENOM has publicly shown one thing, an agent controlling a fleet-standard F-16 in flight. Multi-ship tactics on those jets remain a stated goal.
The evidence each stage owes
The last step removes the pilot, and the claims there currently come from vendors. General Atomics said in February 2026 that its YFQ-42A flew for more than four hours on Collins Aerospace’s Sidekick software, executing commands sent by an operator on the ground. In July, Air & Space Forces Magazine reported that Anduril’s YFQ-44A fired an AIM-120 at a digital target after a human operator commanded the shot. Both events show an uncrewed jet carrying out assigned tasks under supervision. Neither is a published account of an aircraft making tactical choices against a live opponent.
Readers weighing these announcements should ask stage-specific questions. Of a simulation result, ask how the sensors and the opponent were modeled, and how much performance dropped when the same agent first flew. Of a surrogate flight, ask how many times the safety pilot disengaged the agent and why, and whether the target’s position came from the jet’s own sensors or a data link. Of an uncrewed flight, ask which decisions the aircraft made itself and what it did when the link to its operator degraded.
The Air Force and DARPA should publish that information in unclassified summary for each campaign: disengagement counts, the source of targeting data, and the measured gap between simulated and flown performance. The 2023 dogfights were announced without a scorecard. As agents move from Edwards and Eglin onto aircraft with no cockpit, a no-takeover statistic and a vendor press release are too thin a basis for the officers who will be asked to fly alongside the result.
This analysis draws on the public sources linked in the text. Send corrections to info@defenseautonomyreview.com.


