How Social Media Platforms Use Your Data to Train AI

Most social media apps train AI on your data by default. Here's what each platform does and how to opt out where possible.

How Social Media Platforms Use Your Data to Train AI

Recently, the Amazon-owned Twitch made headlines for announcing that it would allow users to opt out of having their streams, VODs, and chats fed into an AI training machine. This, understandably, made a lot of people upset. But it turns out, Twitch might be one of the more moderate social media sites on this front. Most social media apps are training some kind of AI on your data, and few make it as easy as a single toggle to opt out at all.

Almost every social media site out there makes it extremely hard to know exactly how your data is used these days. Training generative AI models like Google’s Gemini is often lumped in with more mundane (but still machine learning-powered) features like YouTube’s algorithm. Sifting through privacy policies to even find out whether a social app contributes to the kind of generative AI that has proven so controversial is an undertaking that would deter most lawyers.

Even if you find out how your data is being used, many sites simply don’t offer you the ability to opt out. Of all the platforms reviewed here, not a single one that trains AI offers an opt-in model. Put simply, if you logged on today, tech companies took that as your consent to be part of their training data. If there is a way to opt out of even some AI training from a social media site, you’ll find instructions in the list below, sorted alphabetically.

Bluesky

What training do they do? Officially, Bluesky doesn’t train any AI models—although it does employ AI in its development and builds AI features. However, since Bluesky uses the decentralized AT Protocol, your posts (and a lot of other data, like who you block) are public. This means it’s possible for just about anyone to scrape your posts and use them to train AI.

What can you do to stop it? You don’t need to do anything; Bluesky itself doesn’t use your posts to train AI. However, if you’re worried about a third party scraping your data, you may want to think carefully about what you post publicly.

Facebook

What training do they do? It might be easier to ask what data Facebook doesn’t use for training AI. After spending enormous sums on a metaverse that never materialized, Facebook’s parent company Meta is pivoting to AI and rushing to catch up to companies like OpenAI, Anthropic, and Google. The company’s policy on training AI with your data is extremely broad, not only encompassing your posts, photos, and interactions on Facebook, but also data collected from third-party brokers and “information that is available on the internet.” That phrase is so all-encompassing that it’s hard to imagine any data Facebook is technically capable of collecting that it would refrain from using. Facebook says it stops just short of training AI on private messages with friends or family, “unless you or someone in the chat chooses to share those messages with our AIs.”

What can you do to stop it? Unfortunately, unless you’re in the EU, Facebook makes it impossible to opt out of AI training with your data. The company carves out a narrow exception—where it is legally obligated to do so—for the specific scenario where you find personally identifying information about you included in a response from one of its AI tools. If that happens, you can submit a complaint, which Facebook will then review to decide whether to act on.

Keep in mind that Meta considers interacting with Facebook’s AI tools as consent to train future AI models on those interactions. So even trying to find out whether your personal information has been collected could expose you further. In general, the only winning move with Facebook may be not to play. It’s worth noting, though, that even deleting your Facebook account won’t necessarily stop Facebook from training AI with your data—it will only prevent them from collecting new data from your account directly.

Instagram

What training do they do? Instagram is owned by Meta, so many of its policies mirror those covered under Facebook above. That said, Instagram has some specific issues of its own, such as the Muse feature, which briefly let users create AI-generated images of other users without their consent, before quickly removing the feature after widespread backlash.

What can you do to stop it? Similar to Facebook, you can file an objection if you find your private information included in AI responses, and if you’re in the EU, you can opt out of training entirely. Otherwise, your only real options are to delete your Instagram account or at least make it private.

LinkedIn

What training do they do? LinkedIn is owned by Microsoft, which is heavily invested in AI, so you should be cautious about the site using your data for AI training. According to LinkedIn’s official policy, users’ posts, comments, profile data, resumes, and group activity, among other data types, can all be used to train AI models. It’s also worth noting that LinkedIn was sued last year over allegations that it was training AI on private direct messages—an allegation LinkedIn denies.

What can you do to stop it? While LinkedIn uses the same default-on model most companies employ, the site at least allows you to revoke your permission. On the site, head to Settings and Privacy > Data Privacy > How LinkedIn uses your data > Data for Generative AI Improvement. Here, you’ll find a toggle labeled “Use my data for training content creation AI models” that is on by default—switch it off. Note that this only applies to data collected on LinkedIn, not the broader Microsoft ecosystem.

Reddit

What training do they do? Reddit is complicated: the company doesn’t train its own AI models, but it does have deals with both OpenAI and Google to train their models on its data. Reddit has reportedly considered ending those deals, at least in part because the company has seen a drop in traffic as AI encroaches on search. But AI models were training on Reddit data long before any official deals were in place, so ending a formal partnership may do little to prevent AI companies from scraping publicly available posts.

What can you do to stop it? Short of deleting your Reddit account and never posting on the site, there’s nothing you can do. Reddit doesn’t currently offer any tool to opt out of AI training, and even if it did, AI companies have been collecting publicly available posts for years regardless.

Snapchat

What training do they do? Like most social media apps, Snapchat uses your public images, video, and audio to train its generative AI models. The company says it uses this data to develop features like AI Snaps and AI Lenses. Snapchat also briefly had a deal to integrate Perplexity’s AI tools into its search, although that deal was “amicably ended” before a broad rollout, so it’s unclear whether any user data was ever shared with or used to train Perplexity.

What can you do to stop it? Unlike most social media apps, Snapchat’s opt-out is relatively straightforward, if a little buried. Inside the app, open Settings and under Privacy Controls, tap Generative AI Settings. Disable the toggle that says “Allow Use of Public Content.” As always, it’s still possible for third parties to scrape publicly available data, though Snapchat’s design makes that somewhat more difficult. For extra protection, avoid leaving any persistent, public content on your profile.

Threads

What training do they do? Threads is owned by Meta, so it shares the same data training policies as Facebook. Unless required by law, publicly available data on Threads will be used to train AI regardless of your consent, and there is currently no opt-out available to most users outside the EU.

Each platform’s approach varies widely, but the common thread is clear: by default, your content is being used to train AI the moment you post it. Where an opt-out exists, it’s worth taking a few minutes to use it—and where it doesn’t, understanding that reality can help you make more informed decisions about what you share and where.