Aidan Gomez

@aidangomez

Synthetic data derived from production user data of consumer AI tools is used for training. I’ve heard this rumour from both large labs’ employees. In particular, if you’re doing something “interesting” like working on complex math/business/software/bio problems you’re dramatically more likely to get trained on because they filter/up-weight towards those usecases where the model has the most to learn. Even in ZDR and “we won’t train on you” regimes, derivative data is usually carved out. The promise is only not to train on exactly the data you put in, rewritten data is fair game.
打开原帖#511482
  1. Research

    OpenAI's Chief Scientist Issues a Rare Warning: We Grew an Alien Mind, and Our Ability to Monitor It Is Fading
  2. Industry

    Alexandr Wang: if this were an mma match, only one guy would be left standing (hint…
  3. Industry

    Alexandr Wang: 3/ Muse operates with the principle of least privilege, so you can de…