• 4 Posts
  • 11 Comments
Joined 1 year ago
cake
Cake day: June 15th, 2023

help-circle

  • yokonzo@lemmy.worldOPtoTechnology@lemmy.worldGenerative AI will eventually poison itself
    link
    fedilink
    English
    arrow-up
    10
    arrow-down
    2
    ·
    edit-2
    8 months ago

    Tbh I think you’re making a lot of assumptions and ignoring the point of this paper. The small model was used to quickly show proof of generative degradation over itérations when the model was trained on its own output data. The use of opt125 was used precisely due to its small size so they could demonstrate this phenomenon in less iterations. The point still stands that this shows that data poisoning exists, and just because a Model is much bigger doesn’t make sense that it would be immune to this effect, just that it will take longer. I suspect that with companies continually scraping the web and sources for data, like reddit which this article mentions has struck a deal with Google to allow their models to train off of, this process will not in fact take too long as more and more of reddit posts become AI generated in itself.

    I think it’s a fallacy to assume that a giant model is therefore “higher quality” and resistant to data poisoning