Data Strategy for AI: Why Data Quality, Operationalisation, and Compute Power are More Important than Data Quantity
More data doesn't automatically lead to better AI. In this guide, we bring together the core principles of a strong data strategy: why data quality is more important than data quantity, what 'good data' actually is, how to operationalise metrics, why high-quality data is becoming scarcer, and why compute power is also becoming a strategic asset.

Data and AI are becoming increasingly central to almost every organisation. Yet, a persistent misconception exists: that more data automatically leads to better results. In practice, the opposite is often true. It is not the quantity of data that determines the value of your AI, but its quality, structure, and the way you define, organise, and present that data. In this article, we bring together the most important principles: why data quality is more important than data quantity, what 'good data' actually is, how to operationalise a data strategy, why a scarcity of high-quality data is looming, and why compute power is also becoming a strategic asset.
Why data quality is more important than data quantity
In recent years, collecting as much data as possible was the norm for companies. The idea was simple: the more data, the better. But this 'collection frenzy' has resulted in many companies now having mountains of data stored without knowing exactly what that data is worth.
Compare it to an overflowing attic or storage room where all sorts of things have been stored over the years: old toys, broken appliances, forgotten photo albums. Things you don't immediately need, but keep 'just in case'. Only when you have to tidy up, for example when moving house, does it become clear how much unnecessary stuff is there. The same applies to many companies in their cloud environment: there is a 'mountain' of information stored, often unused and unorganised, without anyone knowing exactly what it means. This approach leads to high costs, both financial and ecological, because storage capacity is scarce and expensive. And worse still: this unstructured data often adds little to no value.
Blindly adding enormous datasets also backfires when training AI systems. Just as a human cannot learn all of world history at once, adding too much data can actually lead to confusion, corrupted results, and a loss of focus. Instead of thinking in terms of quantity, it is more important to focus on qualitative data: data that is valuable and reliable, and that provides insight into important trends and performance. It's about collecting data that is:
- Consistent: data that is maintained in the same way over time, making trends more visible.
- Shows variation over time: with data consistently collected over different time periods, you gain insight into changes and developments.
- Answers the right questions: by focusing on data that truly reflects your performance and goals, you get the insights you need for strategic decisions.
Particularly granular data, collected consistently in detail over longer periods, provides companies with deep insight into the variations and patterns in their market. A company that collects accurate and consistent data on customer preferences and purchasing behaviour over time can better respond to changing market dynamics and customer needs. This enables organisations to react more quickly and targeted, which has a direct impact on customer satisfaction and business results.
From collection frenzy to data quality
Do you want your company to transition from collecting data to utilising data? Here are some concrete steps:
- Determine which data is truly valuable. Identify which data genuinely contributes to understanding your company's performance and focus only on that. Ask yourself: which metrics are essential for our success?
- Ensure consistency and quality. Check and improve the quality of the data you already have. Is the data correct and consistently maintained over time? Are there errors or gaps?
- Structure and clean up. Organise your data environment so that all data is easily discoverable and provides immediate insight. Use clear structures and ensure everything is well-documented.
- Consider sustainable storage. By storing only qualitative data, you save storage space and reduce costs. This not only helps your business but also contributes to sustainability by reducing pressure on data centres.
What exactly is 'good data'?
Good data forms the basis for reliable AI insights and smart decision-making. But good data goes beyond simply 'a lot of data'. It's about data with specific characteristics that make it useful for AI applications. Good data is:
- Accurate: data without errors or inconsistencies, meaning correct and current information without inaccuracies.
- Complete: the data provides a full picture of the situation you want to analyse, without gaps or missing values.
- Relevant: only collect data that contributes to your specific AI objective and avoid superfluous data that muddles your analyses.
- Consistent: data must be collected and stored in a uniform manner. Variations in measurement units, terminology, or data formats disrupt AI processes.
Practical steps to ensure good data
- Organise and label data correctly. Ensure data is clearly and consistently organised, and apply labels (e.g., 'customer data' or 'transactions') so your AI system can interpret and use this data efficiently. If necessary, use data storage tools that automatically assign labels and metadata.
- Conduct regular data audits. Periodically check the quality of your data to identify inaccuracies and make improvements. For example, schedule monthly audits to ensure both data quality and relevance.
- Jointly define important metrics. Determine with your team which metrics and data points are important. Are you measuring customer loyalty? Then decide whether this is based on purchase frequency, visit duration, or other variables, and clearly document each metric so everyone uses the same definitions.
Give context and meaning to data
Data only becomes valuable when you give it meaning. Numbers and facts without context are useless for AI. Take a customer profile: data can show that a customer buys regularly, but it's the labels and context that indicate why this happens and what it means for your objectives. Therefore, provide metadata that adds extra information to your datasets, so the AI system knows how data relates to each other and what the specific context is.
The role of data governance
Data governance is a system of policies, rules, and procedures that ensure your data remains accurate, secure, and usable. Without good data governance, you run the risk of inconsistent and unreliable analyses. Therefore, appoint a data governance team responsible for maintaining quality standards and compliance. Furthermore, avoid common data errors such as inconsistencies, missing values, and lack of standardisation: consistently use the same data notations and units of measurement, and minimise errors through standardised data collection processes.
Operationalisation: are you really measuring what you want to measure?
An important but often underestimated element in a data strategy is operationalisation: translating abstract concepts and objectives into concrete, measurable quantities. Operationalisation goes beyond simply 'measuring something'; it's about carefully determining what you measure, how you measure it, and why precisely that data forms the core of your decision-making process.
Take customer loyalty as an example. One team might define this as: "A loyal customer visits our website at least three times a week." Another team within the same organisation, however, might describe customer loyalty as: "A loyal customer makes at least one purchase above 50 euros per year." This creates duplicate definitions and inconsistent measurements, with all the consequences that entail. When each department uses its own interpretation, it becomes harder to draw reliable conclusions: one dataset says your customer base is very loyal, while another dataset denies this. This inconsistency can lead to miscommunication, inefficient decision-making, and ultimately strategic failures.
Operationalisation does not stop at the definition alone. Equally important is the choice of data sources. Which data truly reflect your metric accurately? Is it transaction data, website visit duration, repeat purchases, or the average order value? The choice of the right data sources requires precision.
To properly approach operationalisation, these guidelines help:
- Set clear goals. First, define exactly which question you want to answer.
- Create internal consensus. Work with different departments to arrive at a single, clear, organisation-wide accepted definition.
- Document everything. Record all definitions and chosen data sources, so everyone speaks the same language and uses the same starting points.
- Evaluate and improve regularly. The market and customer needs change, so periodically review definitions and measurement methods.
Models can only deliver value if they work with clear, unambiguous data. If your organisation knows exactly what is being measured and why, then the insights from analyses and AI applications will be more reliable and useful. Well-thought-out operationalisation is thus the foundation for everything you do with data and AI.
Provide data incrementally: the 'chain of thought' approach
Not only the quality of data matters, but also the way you present it to an AI model. Just as a human cannot learn all of world history at once, it is much more effective for a language model to be presented with information in a structured, step-by-step manner. This process is known as the 'chain of thought' approach: a method where you provide a model not only with data but also with structured reasoning and context, guiding it step-by-step through the instructions and data it needs.
A practical example: suppose you want to train a model to write texts about history. Instead of offering all historical eras at once, you might start with the Middle Ages, then move on to the Roman era, and so on. Through this phased approach, the model learns both the facts and the connections between different eras. This helps the model to:
- Generate relevant output: the model knows what it needs to do and stays focused on the task.
- Improve quality: by providing context and examples step-by-step, you prevent errors and unnecessary noise.
- Remain flexible: the model can better handle new, unexpected input if it has learned a clear 'thought process'.
If you want to apply this within your organisation, start with clear goals (what tasks should the model perform, e.g., customer service, text generation, or data analysis), provide structured and relevant data, use example instructions that show how the model should approach tasks, and regularly evaluate and optimise the output.
Data scarcity and the trap of synthetic data
About 10 years ago, we spoke about 'big data' and the enormous amounts of data becoming available, and how we needed to 'monetise' them. Since then, the quantities of data have exploded: in the last two years, we have collected 90% of the total amount of data, according to Statista. This increase is partly due to our internet consumption behaviour and partly due to the increase in compute power, which increases the demand for data and generates new data. However, this has led to a major problem: the quality of data is significantly decreasing, and high-quality data is becoming increasingly scarce. Epoch.ai even predicts that we will have 'consumed' all available data by 2026.
Compare it to a gas supply: if more gas is consumed than produced, scarcity arises. The same is now happening with data, with an important nuance: there is sufficient data, but there is a lack of high-quality and usable data.
There are two main reasons for this scarcity. Firstly, due to the enormous increase in AI and language models, the demand for data has grown exponentially; data is, after all, the fuel for AI. Secondly, the rise of synthetic data has exacerbated the situation. Synthetic data is AI-created or derived data, such as AI-generated images or texts. These are often used as training data for AI models, but that creates a vicious circle: if a language model gives an incorrect answer, that output can still be used for (re)training purposes, which can further reduce the quality of data and models.
Consequently, there is enormous demand for datasets with unique, high-quality data, especially data collected directly based on human behaviour in the physical world. An example is the extensive photo and film archive of the BBC, which has been approached by tech companies for access to millions of never-broadcast recordings. These images and sound recordings are sorely needed for the further development of AI models such as image generators DALL-E and Midjourney, and for training models to recognise specific objects. Another example is the multi-million collaboration between Google and Universal Music to gain access to all sound recordings and the rights to use them, aimed at high-quality input for, for example, speech recognition. Companies that collect unique data will be able to earn a lot of money in the coming years by selling it.
High-quality data is also essential to prevent biases. Biases arise when the data used to train AI contains prejudices. These prejudices can be reflected in the AI results, leading to undesirable and discriminatory outcomes. By using diverse and representative data, biases can be minimised as much as possible.
Compute power as a strategic asset
Besides data, the compute power behind AI and advanced algorithms is also becoming a strategic asset. Where data used to be primarily stored and processed in on-premise data centres or via cloud migrations, a shift is now underway. The giants of Silicon Valley, from Sam Altman to Elon Musk and Jensen Huang of NVIDIA, all realise how decisive compute power is for the future of technology and business operations.
A recent example is Elon Musk, who launched Colossus, one of the largest compute clusters in the world. This is used for training his own foundational model, Grok, under his company X.AI. Reportedly, Colossus runs on more than 1,000 NVIDIA H100 GPUs: chips specifically designed to process enormous amounts of data for AI models.
For many companies, compute power has barely been a strategic consideration until now; the focus was on data storage via central servers and later cloud solutions. But with the rise of generative AI, this playing field is drastically changing. GPUs (Graphics Processing Units) and LPUs (Language Processing Units) are becoming the engines of the new economy. Where these chips were once used for gaming or scientific research, they now drive the AI revolution, from complex image processing to training language models and voice interaction.
The challenge is access to that compute power. Especially in Europe, the strategic procurement of this advanced hardware is not a given: scarcity and high costs make access difficult. Therefore, companies increasingly rely on cloud computing from giants such as AWS, Google Vertex and Microsoft Azure, or newcomers like Groq. But there's a downside: as more companies adopt AI, the demand for compute power increases, which drives up costs and can ultimately become a limiting factor for innovation.
The workplace of the future, then, is no longer about human brainpower, but about the ability to feed AI algorithms with sufficient compute power. This leads to a new reality where the workplace is more likely to be a handful of hyperspecialists directing an army of AI bots than a busy office with hundreds of employees. Europe is lagging behind in this: many companies still lack a clear strategy for purchasing or gaining access to compute power. The question, therefore, is: what is your business strategy regarding compute power? Now is the time to develop a vision in this area with your management team and Board of Directors.
Conclusion
More data isn't always better. The time when collecting as much data as possible was the norm is over; nowadays, it's about data quality, consistency, and the ability to derive valuable insights from what you have. This starts with accurate, complete, relevant, and consistent data, with unambiguous definitions that you meticulously operationalise, and with a step-by-step, structured way of presenting it to your models. At the same time, the scarcity of high-quality data is increasing, and access to compute power is becoming a decisive success factor. Companies that have their data and their compute power in order lay the solid foundation for AI that truly delivers results.
- Understanding AI Token Costs

Job van den Berg is an AI keynote speaker, tech entrepreneur and author of five books on AI. He ships AI agents into production every week and delivers 150+ keynotes a year on AI agents and agentic commerce.
On EditieNL I discussed whether AI could threaten humanity. About Anthropic researcher Evan Hubinger, agentic AI, the black box, and why we are building faster than we understand.
A simple AI video already costs about 4 litres of water and as much electricity as a 10-watt LED lamp burning for 42 hours. What happens with full films and commercials, and why digital is not automatically sustainable.
AI keeps getting better, yet workplace sentiment about AI is deteriorating. Research shows why adoption is as much a social challenge as a technological one: from the Matthew effect to psychological safety.
Human in the loop sounds reassuring. But researchers warn that prolonged use of autonomous AI systems can undermine the cognitive capacities of the very supervisors we depend on. The question is not whether a human is formally present, but whether that human can still meaningfully intervene.










































