In the AI Era, the Data Signal Still Comes First

What conversational AI, advertising optimization and marketing mix modeling taught me about the foundation of AI transformation.

Since the rapid growth of large language models, almost every industry has been discussing AI transformation.

Organizations are exploring copilots, AI agents, automated analysis and conversational interfaces. Marketing is no exception. Agencies, platforms and internal teams are all looking for ways to use AI to accelerate work and improve decisions.

LLMs are genuinely useful. During my career break, I have been applying techniques such as retrieval-augmented generation, tool calling and chat-completion APIs to marketing analytics proof-of-concept projects.

However, I do not believe that implementing an AI model should automatically be the first step in a transformation programme.

AI strategy can begin immediately, but implementation should not move ahead of the data signals required by the use case.

Across my experience in banking conversational AI, digital marketing and marketing mix modeling, the technology has changed substantially. The underlying dependency has not:

A stronger model cannot recover a business truth that was never defined, captured or supplied correctly.

“Garbage in, garbage out” still applies—but in the AI era, the meaning of data is broader than a training dataset.

It includes the context used to generate an answer, the feedback used to improve a system, the conversion event an advertising algorithm optimizes towards, the media and outcome data used in a model, and the measurements used to determine whether an AI initiative created value.

AI does not remove the need for reliable signals. It increases the speed and scale at which those signals influence decisions.

When the signal is strong, AI can multiply its value. When the signal is weak, AI can multiply the error.


What do I mean by a data signal?

A data signal is the information that helps a system understand what is happening, what good performance looks like and whether its output was useful.

Different AI and analytics applications depend on different types of signals.

Context signals

These provide the information a system needs to answer or act appropriately. Examples include product documentation, internal policies, customer history and business definitions.

Feedback signals

These indicate whether previous outputs were correct, useful or appropriate. Conversation reviews, annotations, user ratings and expert evaluations are examples.

Optimization signals

These tell an algorithm which outcome to maximize. In digital advertising, this might be a verified purchase, qualified lead or successful application.

Analytical and measurement signals

These support statistical analysis and show whether a business outcome has changed. Media spend, sales, conversion volume, pricing, seasonality and campaign timing can all play this role.

The exact signal differs by use case, but the dependency remains the same: if the signal does not represent the business reality accurately, the system cannot reliably optimize, explain or learn from it.


Conversational AI improved through real customer signals

Earlier in my career, I worked on Stacy, Standard Chartered Bank Hong Kong’s enterprise virtual assistant.

Stacy used NLP models and contextual-vector representations to understand customer enquiries across different banking intents. It could answer general questions and connect to backend services. For example, when a customer asked for a nearby ATM or branch, the system could call the relevant location API and return the results.

The underlying AI model was important, but delivering an enterprise chatbot required much more than a model.

A significant amount of effort went into the infrastructure and operating processes around it:

  • collecting and preparing the required knowledge and training data;
  • integrating the chatbot with digital channels and backend services;
  • establishing model-update and production-deployment processes;
  • monitoring live performance;
  • reviewing conversations across the full range of FAQ intents;
  • regularly annotating customer enquiries; and
  • feeding quality data and operational insights back into model improvement.

After Stacy launched, the actual conversations between customers and the chatbot became an important source of feedback.

Customers did not always ask questions using the phrases anticipated during initial training. They used abbreviations, incomplete sentences, spelling variations and new forms of expression. New enquiry patterns also appeared as customer needs and products changed.

Regular annotation and review helped translate these real conversations into data the model and team could use.

The operating loop was:

Real customer conversations → quality review and annotation → performance insights → model and journey improvement → new conversations

Insights generated through this process contributed to reducing the chatbot’s negative-feedback rate from 7.4% to 4.1% and lifting year-over-year cost savings by 27.5%.

Those results did not come from data alone. The model, infrastructure, vendor delivery, product design and operating team all mattered.

The lesson is that the model by itself was not sufficient. Sustainable improvement required a feedback system that could capture real customer behaviour, turn it into quality data and use that data to improve the service over time.

Modern LLMs are much more flexible than the NLP systems used in that period. They are less dependent on a narrowly defined list of phrases and can handle open-ended language more effectively.

But the need for feedback has not disappeared.

An organization still needs to know:

  • whether answers are correct;
  • whether they resolve the user’s need;
  • where failures occur;
  • which new questions are emerging;
  • when human escalation is required; and
  • whether the system is generating the intended business return.

A more capable model may improve the starting point. It does not replace the operating data loop.


Advertising AI depends on the conversion signal

The same principle appears in digital advertising, although the data serves a different purpose.

Google, Meta and other advertising platforms use sophisticated machine-learning models to optimize campaign delivery. These models can adjust bids, identify likely converters and allocate budget across placements at a scale that would be difficult to manage manually.

However, an advertising algorithm cannot optimize towards the correct outcome if that outcome is not accurately captured.

Consider a few common measurement problems:

  • A purchase event fires twice for one transaction.
  • A lead conversion records a button click rather than a successful form submission.
  • Test transactions are mixed with genuine customer activity.
  • Important event parameters are missing or inconsistent.
  • Different platforms use conflicting definitions of the same conversion.
  • Consent and browser restrictions reduce the available signal.
  • A low-value action is treated as being as important as a completed sale.

The platform may still optimize successfully—but it will optimize towards the signal it receives, not necessarily towards the business outcome the marketer intended.

This is why foundational MarTech work remains vital even when advertising platforms have powerful AI.

That work includes:

Measurement design

The business must define what a conversion actually means. A button click, form submission, qualified lead, approved application and completed sale are not interchangeable outcomes.

Tag setup and quality assurance

Tags, triggers and event parameters must be configured and tested so that the required actions are captured consistently without duplication.

Datalayer enhancement

A structured dataLayer can provide reliable information about user actions, transaction details and business context rather than relying on fragile page elements or inferred behaviour.

GA4 implementation audits

Regular audits can identify missing events, incorrect parameters, duplicate tracking, inconsistent definitions and implementation gaps.

Server-side tagging and conversion APIs

Server-side tagging and platform conversion APIs can strengthen first-party signal collection and provide greater control over how data is validated and transmitted, subject to consent, privacy and platform requirements.

These projects may receive less attention than a new AI agent, but they are part of the data infrastructure that advertising AI relies upon.

The advantage does not simply come from sending more events. It comes from sending signals that more accurately represent valuable business outcomes.

A clean, verified signal for a successful application can be more useful than a large volume of noisy engagement events.

Before asking whether an advertising platform’s AI is performing well, marketers should first ask:

Are we giving the algorithm an accurate definition of success?


In MMM, data preparation is often the real bottleneck

The same dependency appears in advanced analytics such as marketing mix modeling.

MMM attempts to estimate how marketing activity and other business factors relate to an outcome such as sales or conversions. The modeling engine may be sophisticated, but the reliability and interpretability of the results still depend heavily on the input data.

In practice, running the model is often not the most time-consuming part of the project.

A substantial amount of work may be required before modeling can begin:

  • collecting spend and activity data from different markets or platforms;
  • resolving discrepancies between data sources;
  • aligning reporting periods and campaign flighting dates;
  • standardizing inconsistent channel names and taxonomies;
  • identifying missing offline media activity;
  • reconciling sales or conversion definitions;
  • adjusting for changes in data collection over time; and
  • finding defensible alternative or proxy data when the preferred source is unavailable.

These are not minor technical details. They affect what the model can learn from the data.

For example, if Brand Search and Generic Search are combined into a single channel, the model may have difficulty distinguishing their different roles. If media activity is recorded in inconsistent time periods, the relationship with the outcome becomes harder to interpret. If an important offline channel is missing, the model may attribute part of its effect elsewhere.

Missing, inconsistent or poorly grouped media data can weaken the model’s ability to distinguish channel effects, increase uncertainty and make the results more difficult to interpret reliably.

This is why data collection, reconciliation and taxonomy design can take longer than the modeling itself. The team may need to rectify the source data, find an alternative source or construct and document an appropriate proxy before proceeding.


GenAI can accelerate MMM—but only after the data foundation exists

Before my departure from WPP Media, I supervised the development of an internal GenAI-enabled MMM application.

The proposed workflow was designed to make the modeling process more accessible:

  1. The user uploads a cleaned marketing dataset.
  2. The user identifies the relevant input columns.
  3. The user selects a sales or non-sales conversion outcome.
  4. The application runs Google Meridian.
  5. GenAI produces an initial summary of the MMM results and potential insights.

This can reduce some of the manual work required to operate the model and prepare an initial interpretation. It can also make the output easier for a wider group of users to access.

But the first step remains important: the user uploads a cleaned dataset.

The GenAI layer does not resolve an incorrect channel taxonomy, recover missing historical spend or determine automatically whether a proxy is methodologically defensible. It can summarize the model results, but the quality of that summary still depends on the quality of the analysis underneath it.

A polished AI-generated explanation cannot compensate for unreliable input data.

This does not reduce the value of GenAI. It clarifies where GenAI creates value.

Once the data has been prepared and the model has been specified appropriately, GenAI can help:

  • explain technical outputs in more accessible language;
  • draft an initial summary;
  • identify results that may require further investigation;
  • organize findings for different stakeholders; and
  • reduce repetitive reporting work.

The right role for GenAI is to accelerate access to and interpretation of a sound analysis—not to give weak analysis a more convincing presentation.


The AI value stack

My experience across these areas can be summarized through a simple dependency stack:

Data → Analytics → Advanced Analytics → AI

Each layer creates capabilities that the next layer can use.

1. Data: Capture the signal

This layer establishes the operational foundation:

  • event and conversion definitions;
  • tag setup and testing;
  • dataLayer design;
  • customer conversation data;
  • CRM and transaction data;
  • first-party data collection;
  • data pipelines;
  • taxonomy and metadata; and
  • data quality and governance.

The objective is to capture business activity accurately and consistently.

2. Analytics: Establish a shared view of performance

Analytics turns collected data into understandable information through:

  • reporting;
  • dashboards;
  • KPI definitions;
  • segmentation;
  • diagnostic analysis;
  • quality monitoring; and
  • performance measurement.

This layer helps an organization establish what happened and whether stakeholders agree on the definitions.

3. Advanced analytics: Estimate, predict and explain

Advanced analytics applies statistical, econometric and machine-learning methods to support more complex questions:

  • marketing mix modeling;
  • propensity modeling;
  • forecasting;
  • clustering;
  • attribution analysis;
  • experimentation; and
  • causal measurement.

These methods rely on the definitions, history and data quality established below them.

4. AI: Accelerate access, interpretation and action

AI can make the capabilities underneath more accessible and scalable through:

  • conversational interfaces;
  • code generation;
  • automated first-draft analysis;
  • retrieval-grounded assistance;
  • generated summaries;
  • tool-using workflows; and
  • bounded AI agents.

This should not be interpreted as a rigid maturity model in which an organization must complete every layer before experimenting with AI.

AI can also assist with earlier layers. It can draft data-quality code, propose event taxonomies, review tracking requirements or summarize analytics documentation.

But an AI application cannot reliably use a business signal that has never been defined, captured or governed.

The layers can develop together, but the dependencies remain.


AI readiness is use-case specific

No organization has perfectly clean data, and waiting for perfection can become another reason not to innovate.

The goal is not to fix every dataset before starting any AI project. The goal is to identify the signals required by each use case and determine whether they are fit for purpose.

A general writing assistant may not require access to an organization’s customer data. An advertising-optimization project requires a reliable conversion signal. A customer-service assistant needs approved knowledge, conversation feedback and escalation data. An MMM application needs consistent media and outcome data.

AI readiness should therefore be assessed at the use-case level.

Before implementation, I would ask:

What business process or decision are we improving?

The initiative should begin with a real business requirement, not only a desire to use a particular model.

What signal tells the system what good looks like?

Is it customer feedback, a verified conversion, revenue, an expert label or another measurable outcome?

Is that signal defined and captured accurately?

Different teams may use the same term—such as “lead,” “conversion” or “engagement”—to mean different things.

What context does the system need?

The answer may depend on internal knowledge, customer history, product rules, source documentation or other business data.

How will the signal be monitored?

A tracking implementation, taxonomy or customer behaviour can change over time. Data quality must be monitored after launch.

What feedback will improve the system?

A production AI application needs a way to learn where it performs well, where it fails and when users need human support.

What happens when the data is incomplete or wrong?

The system may need validation rules, confidence thresholds, warnings, fallbacks or expert escalation.

What business return should the initiative create?

The expected benefit should connect to revenue generation, cost reduction, productivity or competitive advantage.

These questions connect the data foundation to the broader translation problem in AI adoption.

Translating a business requirement is not only about explaining it to a technical team or turning it into a prompt. It also means translating the desired business outcome into a signal the system can observe and use.


Final thought

The current excitement around LLMs and AI agents is justified. These technologies can accelerate work, make complex capabilities more accessible and create new ways for people to interact with data and software.

But they are not shortcuts around unresolved data problems.

Across conversational AI, digital advertising and marketing mix modeling, I have seen data play different roles:

  • Real customer conversations and annotations provided feedback for improving a chatbot.
  • First-party conversion events told advertising algorithms what to optimize.
  • Media and outcome data determined what an MMM could estimate and explain.
  • Cleaned modeling outputs gave GenAI something meaningful to summarize.

The technology and use cases were different, but the pattern was consistent.

AI does not remove the need for a reliable data signal. It increases the speed and scale at which that signal affects decisions.

This is why AI transformation should not begin only with questions such as:

  • Which model should we use?
  • Which copilot should we launch?
  • Where can we deploy an AI agent?

It should also ask:

What signal will tell the system what is true, what is valuable and whether it is working?

When that signal is strong, AI can become a powerful multiplier.

When it is weak, a stronger model may only produce a faster, more scalable and more convincing version of the wrong answer.

The path to sustainable AI value remains:

Data → Analytics → Advanced Analytics → AI

The tools have changed. The foundation has not.

Leave a Reply

Your email address will not be published. Required fields are marked *