SyncAI.news, a Varaisys broadcasting
​Why Your AI Budget Is Really A Data Cleanup Bill In Disguise
LG

Leon Gordon, Forbes Councils Member

· 1 min read

World NewsForbes: Innovation

​Why Your AI Budget Is Really A Data Cleanup Bill In Disguise

Leon Gordon: CEO & Founder of Onyx Data, 6× Microsoft MVP, building governed enterprise AI and data solutions for business impact.

A board approves seven figures for an AI program. The line item says “AI.” Twelve months later, the model works in a demo, stalls in production and the post-mortem reads like a plumbing report: mismatched keys, stale records, undocumented lineage, a customer table nobody trusts. The money was booked as intelligence. It was spent on remediation.

This is the pattern I see across enterprise estates. An AI program brings an old bill into the light: the data debt an organization has accumulated for years, discovered when a demo has to become production. The model is the visible line item. The condition of the data beneath it decides whether the program survives.

Research and practitioner evidence have pointed in the same direction for a decade. Google’s engineers warned in a 2015 NeurIPS paper that only a small fraction of a real-world machine learning system is the model itself; the surrounding system of data dependencies, configuration, glue code and monitoring carries “massive ongoing maintenance costs.” In a 2021 CHI study of 53 practitioners working on high-stakes AI, “data cascades,” the downstream damage from upstream data problems, showed a 92% prevalence among those studied.

In Anaconda’s 2020 survey of 2,360 data professionals across more than 100 countries, respondents reported spending an average of 45% of their time loading and cleansing data before they could use it to develop models and visualizations.

Three Places The Estate Breaks

The first is contaminated ground truth. A 2021 NeurIPS analysis found at least 3.3% label errors on average across 10 widely used machine-learning benchmarks, including at least 6% of the ImageNet validation set. The authors also showed that past specified levels of label noise, a smaller model can outperform a larger one. You can pay for more capability and get less because the data underneath is wrong.

Original source

This story was published by Forbes: Innovation and written by Leon Gordon, Forbes Councils Member. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on forbes.com

Similar News