OpenAI says it has evidence China’s DeepSeek used its model to train competitor

Posted by IHateTrains123

11 Comments

  1. I have no idea what im talking about article is paywalled but why is that bad? If I make a Windows competitor using a Windows computer no one would care.

  2. Flaky-Ambition5900 on

    The problem is that there is so much ChatGPT generated text on the internet that you really can’t avoid it for model pretraining.

    Bots have been spamming ChatGPT text everywhere so anything that trains on the Internet will be compromised with ChatGPT text. It doesn’t really matter if DeepSeek wanted to train on ChatGPT output or not, ChatGPT output will have made it into DeepSeek’s training data anyways.

    (As a side note, there are interesting theories about how ChatGPT spamming will eventually make it impossible to train good chatbots as the noise will override any remaining signal. It’s very similar in theme to disappearing polymorph stuff. See https://www.nature.com/articles/s41586-024-07566-y)

  3. At this point the AI industry is just hypocrites finger pointing at other hypocrites while they each keep building their glass house with overblown venture capital

  4. moffattron9000 on

    If you spent years stealing anything ever written on the internet ever (including this very lame, snarky comment), you don’t get to blame someone for doing the same thing to you.

  5. So the company that’s being sued for stealing content to train its AI is now mad that someone stole their content to train AI? Love it.

  6. I wouldn’t care about this if Christ himself came down from heaven and told me I have to

Leave A Reply