Data is the Hard Part
ChinaTalk Jordan Schneider
The Pentagon’s AI problem isn’t the model. It’s the data. Without steady collection, labeling, and rules, the clever stuff never gets to troops.
Based on reporting by ChinaTalk, Jordan Schneider — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Bharat Patel’s blunt claim is that “AI-ready data” is mostly a fantasy. Data, he argues, is always messy; what matters is the use case, because the use case decides what data you need and how clean it has to be. That framing runs through the whole conversation: if the problem isn’t pinned down, the data work never really starts.
He points to Project Maven, begun in 2017, as the first big lesson. The team wanted computer vision models for specific targets, but a lot of the data they got didn’t actually contain those targets. The result was predictable: models that looked promising in the lab fell apart in operations. The fix wasn’t magic. It was continuously collecting relevant data and treating that as part of the job, not an afterthought.
The same lesson showed up in Army work on tank sensing. There wasn’t tank data ready to use, so the team built a data collection box and put it next to the sensor they cared about. They started with second-gen FLIR imagery because that’s what they had. But Patel says the real path to autonomy is multimodal data — not one sensor, but several — because that’s what lets systems track targets and understand what they’re seeing.
He is also skeptical that autonomy is close to replacing humans in armored vehicles. A human still has to verify what the machine is doing, and Patel says the nation is not ready to let large numbers of autonomous platforms roam around unchecked. Mistaking a civilian vehicle for an enemy one is too easy, and policy hasn’t caught up.
Ukraine is the clearest example of how this evolves. Patel says what exists there today came after years of collecting video and sensor data, moving it to a central place, and learning from it. In his telling, the U.S. military still lacks that kind of continuous machine-learning pipeline. Too often data gets erased, deleted, or dumped somewhere it loses context. The “boring” pieces — standards, governance, testing, and pipelines — are the real bottleneck, and he thinks the government should keep control of the data while industry does most of the work.
My take — AI-written commentary, not fact-checked reporting
This is the least sexy AI lesson and also the one everybody keeps dodging: if the data plumbing is bad, the model is theater. Defense loves grand language about autonomy, then acts surprised when the spreadsheet and the storage bucket matter more than the demo. There’s a reason bureaucracy keeps winning — it’s where the real war on AI is being fought.
Read more about this at: ChinaTalk