# Anthropic Project Pilot / Drone-Bench source note - 2026-07-27 Purpose: weekly Managing Expectations AI Papers Library maintenance note. Selected item: Anthropic and Andon Labs' July 2026 `Project Pilot: Can AI control a drone?` / Drone-Bench release on AI agents controlling low-cost drone hardware for a simple locate-and-follow task. ## Bottom line Substantive source found and selected for one new library/blog note. Anthropic published `Project Pilot: Can AI control a drone?` on July 24, 2026, in collaboration with Andon Labs. Andon Labs also published `Drone-Bench`, a benchmark measuring whether AI models can write code to control an off-the-shelf drone in a simple indoor surveillance-style locate-and-follow task. Editorial framing: treat this as frontier-lab red-team evidence about physical-agent capability growth and dual-use risk, not as proof that AI systems can reliably operate drones in real-world conditions. The reported experiment used a slow indoor drone, one office floorplan, a limited number of people, consent from the person being followed, and benchmark decomposition into subtasks. Anthropic says current models are still bottlenecked by reconstruction/localization errors and that more realistic tests would be needed to assess operational capability. ## Sources checked this run ### Anthropic Research feed - Page: https://www.anthropic.com/research - Accessed: 2026-07-27 - Tooling: direct HTML retrieval via Python/urllib because general web_search/web_extract was unavailable in this Hermes runtime. - Recent visible items included: - `Project Pilot: Can AI control a drone?` - Jul 24, 2026. - `How Canada uses Claude: Findings from the Anthropic Economic Index` - Jul 14, 2026. - `Claude's values across models and languages` - Jul 13, 2026. - `Claude plays robotics` - Jul 9, 2026. - `An off switch for dual-use knowledge in AI models` - Jul 8, 2026. - `A global workspace in language models` - Jul 6, 2026. - Local duplicate check: Project Fetch phase two, global workspace, and off-switch/GRAM already have local Managing Expectations notes. The Project Pilot / Drone-Bench item had no existing local article/source note. ### Anthropic article / official source - Page: https://www.anthropic.com/research/project-pilot - Accessed: 2026-07-27 - Title metadata: `Project Pilot: Can AI models fly drones?` - Page headline: `Project Pilot: Can AI control a drone?` - Page date: Jul 24, 2026 - Authors/source line: Anthropic and Andon Labs - Page description metadata: `We worked with Andon Labs on Drone-Bench, a new benchmark testing whether AI models can autonomously fly a drone to locate and follow a person.` Key facts verified from page text: - Anthropic frames the work as a continuation of physical-world red-team projects such as Project Vend and Project Fetch. - The task is an indoor office locate-and-follow task using a quad-rotor drone and a reference photo of a consenting person on the experiment team. - Anthropic says drones are dual-use: useful in fields such as agriculture, search and rescue, disaster response, and lawful public-safety work, but also relevant to surveillance, warfare, privacy, and physical-security risks. - The task was decomposed into five subtasks: Reconstruct, Localize, Navigate, Detect, and Follow. - Anthropic says Drone-Bench was created by Andon Labs in consultation with Anthropic; Anthropic says it did not have access to Drone-Bench and that Andon Labs ran the evaluations Anthropic reports. - Anthropic says Andon tested 15 models from three developers, including GPT, Gemini, Opus, and Fable model families. - Anthropic reports that newer models get successively further on the subtasks. Detection and following were easiest; reconstruction and localization were hardest. - The best-performing model in the reported results was Claude Fable 5, which passed the baseline on all tasks except reconstruction. - In a real-drone end-to-end test, Fable 5 performed better than the baseline at detecting and following, but reconstruction errors compounded into localization/navigation errors and the model was unable to autonomously navigate between rooms. - Anthropic says current models reached the human-AI baseline in at least one simulation for four of the five tasks, but even Fable 5 reached the baseline on average for only three of five tasks. - Anthropic's limitation language: drones moved at slow speed; only one office floorplan was tested; there were a limited number of people; Andon did not test outdoors in large crowds; more realistic and diverse experiments would be needed to assess operational capability. - Footnote: the hardware was a DJI Tello EDU, which Anthropic says retailed for $129 at the time; the followed person consented and was a member of the experiment team. ### Andon Labs Drone-Bench page - Page: https://andonlabs.com/evals/drone-bench - Accessed: 2026-07-27 - Title metadata: `Drone-Bench | Andon Labs` - Description metadata: `We're releasing Drone-Bench, a benchmark measuring how well AI models can write code to surveil real-world environments on low-cost drone hardware.` Key facts verified from page text: - Andon describes Drone-Bench as a benchmark for simple drone surveillance capabilities of frontier models. - Andon says the demo uses an off-the-shelf drone autonomously navigating an office to find and follow a specified person. - Andon says the demo spans reconstruction, localization, navigation, detection, and following. - Andon states that each task is scored against a human baseline, specifically demo code written by a human working with coding agents. - Andon reports `84% progress` as mean task progress toward the baseline on the visible benchmark page. - Andon says no AI lab will be able to train on the eval. - Andon notes that each task is run in isolation using clean baseline upstream artifacts so per-task scores are isolated; end-to-end error compounding is discussed separately. ### OpenAI official RSS / blocked direct-page check - RSS: https://openai.com/news/rss.xml - Accessed: 2026-07-27 - RSS was reachable and listed recent July 2026 items, including: - `How AI is expanding what people do at work` - Jul 27, 2026. - `Launching Health in ChatGPT` - Jul 23, 2026. - `Advancing the next era of national science` - Jul 22, 2026. - `Safety and alignment in an era of long-horizon models` - Jul 20, 2026. - Direct OpenAI pages returned HTTP 403 in this runtime, so OpenAI items were treated as official RSS leads only. They were not selected because the full primary article body could not be independently verified during this run. ### Google DeepMind blog / official source - Page: https://deepmind.google/blog/ - Accessed: 2026-07-27 - Recent visible DeepMind items included `Introducing Gemini 3.5 Flash Cyber`, `Accelerating the frontiers of scientific discovery: Google's $40M commitment to the Genesis Mission`, `Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber`, and `Our approach to bioresilience`. - Local duplicate check: `Securing the future of AI agents` / `GDM AI Control Roadmap` already has a Managing Expectations source note and article from 2026-07-06. - Selection rationale: the Anthropic / Andon Labs Drone-Bench item was the clearest new source-grounded technical red-team item for the AI Papers Library this run. ### LawZero / Yoshua Bengio - Homepage: https://lawzero.org/en - Accessed: 2026-07-27 - Recent visible LawZero homepage/news items still included the Scientist AI safety-case publication/news release and ICML 2026 presence material. - Local duplicate check: LawZero/Bengio Scientist AI materials were already captured in `research/ai/yoshua-bengio-update-2026-07-02.md`. ## Why this was selected - Primary source from a frontier AI lab plus the benchmark author's page. - Directly extends the existing Managing Expectations physical-agent lane: Project Fetch / robot dog -> Project Pilot / drone. - It is not merely product marketing: it describes a capability evaluation, benchmark decomposition, model trends, limits, and physical-world safety governance questions. - Strong Managing Expectations angle: AI autonomy leaves the chat window when models can write code for hardware, but capability does not equal safe or reliable operational deployment. ## Article framing used Title: `Anthropic Drone-Bench: AI Agents, Drones and Physical Oversight` Evidence label: `frontier red-team benchmark / physical-agent safety warning` Core caution: Drone-Bench is a serious warning signal about the direction of AI-plus-hardware capabilities, but it does not show reliable autonomous drone operation in realistic public environments. It also should not be read as a how-to guide; the public-interest takeaway is oversight, evidence labels, privacy, and governance. ## Files updated locally - `blog/articles/anthropic-drone-bench-physical-agent-safety.html` - `research/ai/anthropic-project-pilot-drone-bench-source-note-2026-07-27.md` - `ai-library.html` - `ai.html` - `research/ai/ai-paper-library-seed-2026-06-11.json` - `research/ai/ai-paper-library-source-note-2026-06-11.md` - `blog/index.html` - `sitemap.xml`