TechForge

January 16, 2025

  • Ethical AI for marketing to be next niche.
  • Only 12% of 134 major websites respect user preferences for tracking and cookies.
  • ‘Clean data’ to be used in more ethical AI models for marketers.

The use of data in creating end-user profiles is nothing new. Businesses use data, whether collected themselves or obtained through data brokers, to build digital pictures of their customers and prospects.

But despite users’ opt-out preferences set in-browser (such as the ‘do not track’ flag) or expressed during visits to websites, such preferences are ignored in nearly 90% of cases. Those were the findings gathered by an automated sweep of 134 major US websites that used browser emulation to see the extent of unwanted tracking, data collection, and user profiling.

The company behind the exercise, Ketch, estimates that the well-known websites collectively engage in over 1.7 trillion “data collection events” and that 48% of website trackers are concerned with advertising and marketing. Of the remaining tracking cookies encountered, around 14% were unverifiable as to whether or not they were PDTs (privacy-dependent trackers). PDTs should only be placed in a visitor’s cookie jar if the user has accepted them.

Additionally, Ketch found that 38%-40% of PDTs remain after a user opts out. In fact, only 12% of the companies’ websites were fully compliant.

The myth of the ethical AI

The thrust of the paper maintains [email wall] that data so collected – without users’ consent – effectively taints the information that companies use to train AI models, especially those designed to be part of marketing efforts. Conversion metrics (when a user clicks a download link, makes a purchase, and so on) join personal information such as demographic and geographic data, as well as behaviour patterns gathered over a users’ visits to other websites.

As the use of AI becomes effectively ubiquitous, companies may be keen to differentiate themselves from competitors by means of declaring their ethical data gathering practices. It’s into this future niche that Ketch appears to be fitting itself, offering data for marketers that’s distinguished by the fact that its sources agreed to share their information.

If a website owner only wishes to collect consensual data, the simple required step is to respect the visitor’s preferences – something that isn’t happening at present, at least in the majority of cases where large companies’ websites are concerned. However, there remain issues of granularity and directed consent. A site visitor may give their assent to data collection by the site’s owner, but that consent is not transferable to other organisations, unless expressly granted at the time of acceptance. Therefore ‘clean data’ (as opposed to ‘dirty data’ as described by the Ketch report) only has value as a commodity if kept in-house and not given to others.

Given the moral vacuum around user privacy and preferences, it seems likely that data brokers may envisage the emergence of a premium product for monetisation, whether to other commercial organisations or operators of AI models. Feeding information collected in ways at best described as un-consensual into AI models for marketing is the present norm – and continues the process by which AI models are trained.

Attempting to impose a status of virtuousness or higher-quality on information destined for AI ingestion doesn’t somehow extinguish the moral dumpster fire of artificial intelligence. With terabytes of illegally-obtained data forming the basis of all notable AI engines, adding a little consensual seasoning does nothing to remove the stigma. AI models including those built for marketing purposes come pre-trained on huge amounts of non-consented data and therefore are beyond moral rescue.

Ethical AI in the future

Organisations have to decide whether or not they use AI, and if they decide to, be 100% sure of the legitimacy of the data used to train the model they select. That’s the only way to have a legitimate moral stance. The use of ‘clean data’ once the AI is in production is whitewashing.

At present, there are no AIs that have been trained entirely on consented data, although there are several organisations that aim to build them. Given that large model operators are now offering content creators cash for their outtakes, it’s apparent that the next generation of AIs has already run out of training material. Models that can only access data sets deliberately limited by their operators’ ethics may compete against the giants, but it seems unlikely that they would ever be as powerful and thereby win a majority of users.

Author

  • Joe Green

    Joe Green is a writer based in Bristol, UK. He acquired his first Mac and dial-up modem in 1992 and has worked in the tech industry since 2000. He writes and podcasts, specialising in open-source, networking, cybersecurity, software development and online privacy.

    View all posts

About the Author

Joe Green

Joe Green is a writer based in Bristol, UK. He acquired his first Mac and dial-up modem in 1992 and has worked in the tech industry since 2000. He writes and podcasts, specialising in open-source, networking, cybersecurity, software development and online privacy.

Related

August 11, 2026

August 10, 2026

August 5, 2026

July 30, 2026

Join our Community

Subscribe now to get all our premium content and latest tech news delivered straight to your inbox

Popular

12345 view(s)
11326 view(s)
7643 view(s)
6152 view(s)

Subscribe

All our premium content and latest tech news delivered straight to your inbox

This field is for validation purposes and should be left unchanged.
Name(Required)