- Meta’s AI chatbot to be trained on Brits’ data.
- Tech giant will allow opt-out.
- EU still blocking similar activities.
Meta is going ahead with its plans to scrape UK citizens’ data from Facebook and Instagram to train its AI models, with the exception of users’ private messages.
The US-based giant has not had regulatory approval from the UK’s Information Commissioner’s Office (ICO); intstead the ICO will monitor the process after it received concessions from Meta that allow users to opt out from the process.
Since the UK’s departure from the EU, decisions around what is ‘legitimate interest’ in using data are no longer governed by Brussels. The EU has refused to allow its citizens’ data to be used by Meta to train its AI algorithms. The company has accused the EU of preventing development of AI as a technology, stating in a web post, “We believe that Europeans will be ill-served by AI models that are not informed by Europe’s rich cultural, social and historical contributions.”
Meta has confirmed that it will use publicly-shared posts on both platforms to train its models for products like the Meta AI chatbot. “This means that our generative AI models will reflect British culture, history and i diom, and that UK companies and institutions will be able to utilise the latest technology,” the company said.
The fact that Meta is actively seeking local regulatory approval to use citizens’ data is a differentiating factor, it maintains. “Google and OpenAI […] have already used data from European users to train AI. Our approach is more transparent and offers easier controls than many of our industry counterparts already training their models on similar publicly available information.”
The approach taken by Microsoft, Google, Hugging Face and others to train their models has been to scrape the internet, in a manner analogous to search engines that scrape websites. Any accessible information is collated and used to inform nascent machine learning models about the relationships between individual data instances. This has led to controversey around the use of material that is released under specific licences that dictate the terms of its use, copyrighted material and intellectual property. The majority of AI companies have operated – and continue to operate – therefore, under the ‘better to seek forgiveness than permission’ moral decision. Meta’s approach is more circumspect.
It is noteworthy that under the terms of the ICO’s oversight, users will be able to opt-OUT rather than -in to their data being so harvested. Screens that present Terms and Conditions on installation of software are rarely, if ever, actually read and acted on by users, other than to click the ‘Accept’ option and proceed. If Meta’s options for its users’ data to be used in machine learning training are embedded in wider T&Cs, it seems unlikely that many users will opt not to take part. The chances of Meta losing some users along the way if AI training use approval is a secondary end-user query (in a second ‘do you consent’ screen, for example) are higher, but not by much.
Britons will likely be adding their rich vernacular and cultural complexities to the Meta AI chatbot’s model whether they realise it or not. Meta’s engagement with the UK’s ICO is arguably a positive step in the right direction, albeit one that may be guided by Meta’s more cautious lawyers.
Data privacy concerns over large technology companies’ use of personal data remain a lesser concern to most businesses and individuals, who prefer the ubiquity and convenience of platforms like Instagram, Facebook and Google over worries that information about their activities is used and sold. The quiet cries from creative types reacting to the misappropriation of their work in learning algorithms are drowned out by the hoofbeats of the stampede to all things AI.
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is co-located with other leading events including Intelligent Automation Conference, BlockX, Digital Transformation Week, and Cyber Security & Cloud Expo.
Explore other upcoming enterprise technology events and webinars powered by TechForge here.
Author
- View all posts
Joe Green is a writer based in Bristol, UK. He acquired his first Mac and dial-up modem in 1992 and has worked in the tech industry since 2000. He writes and podcasts, specialising in open-source, networking, cybersecurity, software development and online privacy.