Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'm surprised that people here don't care at all about these models openly training on your data, especially if you use them straight from the model developer. Whereas things like "GitHub now automatically opts everyone into using their code for model training" get hundreds of justifiably angry comments, I never see this brought up anymore on posts like these talking about using Chinese models through OpenRouter. This might be explained by "well they're different people", but the difference is very stark for that to be the whole explanation.


The cool thing about open-weights model is that you are free to use alternative providers that won't phone home to the original model creators.

I see 6 alternative providers listed on Openrouter for DeepSeek V4 Pro for example.


At least that’s what they’re telling you. It’s a ”trust me bro” scenario.

I’d rather use the phone home version (deepseeks own endpoint). The benefit is that I’m fairly certain that they actually host the model I’m paying for.


If you're not Chinese, and you start a company outside of China, and your whole pitch is "We run open weights and we have nothing to do with China", 1) why would send data to China?? 2) why would you risk your business to do a thing that makes no sense?


A fly by night operation created primarily for the purpose of collecting training data and corporate espionage will make whatever claims they think will get them the right traffic.


Well, the context was running the models via open router, not hosting 800B> models yourself. Of course, if given the option I believe most people would pick ”don’t share sensitive data”.

What I’m trying to say is that EVERYONE uses your data, even the sensitive type. So you might aswell use an endpoint that does what it says and treat EVERY endpoint whether that’s OpenAI or anthropic as if it’s collecting all of your data.


No, not everyone uses your data. There are providers who very explicitly do not collect or use your data.


Sure, and I won’t collect or otherwise store your credit card info if you send it to me. Trust me bro :)

No but seriously, I am astonished by the level of trust you have for these for-profit companies. I’ll remind you of this quote:

”Zuckerberg: People just submitted it. Zuckerberg: I don't know why. Zuckerberg: They "trust me" Zuckerberg: Dumb fucks”


Some providers are based in the US or EU and would face legal repercussions for lying about what they do with your data. It's a bit more than "trust me bro". Off the top of my head, you can use Fireworks, for example, which is based in California and would face the same consequences for lying about their data policy as OpenAI or Anthropic would.


Meta is based in the US, yet they torrented TERABYTES worth of books to feed their AI.

I’m not trying to be negative here, but your point is invalidated by that particular event in itself.


What, because they broke the law in one way, they'd break the law in every way? That's not how business works. The way business works is, I steal from other people to make a product, but then I don't steal from my customers, because if they find out, then I no longer have any customers. (Plus all their customers would sue them, which would both legally and financially tank them)


That's a naive way of thinking. You're saying "oh, they are thieves in this way, but they surely wouldn't be thieves in this other way!"

If you have no problems shitting on tens of thousands of authors of books, you don't have problems shitting on your customers as well (which they have proved again and again, see https://en.wikipedia.org/wiki/Facebook–Cambridge_Analytica_d...)

Let's just say I wholeheartedly disagree with your viewpoint and leave it there :)


You definitely have a bone to pick. Chinese researchers usually have given the world the most cheap and consistent high quality research around LLMs. They don't pretend, they do the work and release the goodies. Mostly so cheap, every one in the world has a chance to use close to frontier models. Why would you respond with "Anger"?

You let us know what your real complaint is about and let's not feign indignation at open models and research.


You're making completely unfounded assumptions about me. I use Chinese models myself.


Anthropic and OpenAI took your data, trained their model, and tell you "we are not going to tell you anything how we trained our models, we are not giving your the weights our models, you will have to pay us to access the model trained from your data".

they took your rights and your data.

Chinese labs took your data, trained their model, and tell you "this paper details how our models are trained using your data, here is the final weights of our model trained from your data, feel free to use it for what you want, it is your model trained on your data".

they converted your data, everything is still in your hand under your control.

you couldn't see the difference?

Your specific question can actually be translated as -

1. why people don't stop Chinese labs so US monopoly can be maintained?

2. why people don't stop Chinese labs providing free models to those who would otherwise never be able to afford the same $200 USD/month Anthropic and OpenAI subscriptions.

3. why people don't complain Chinese labs publishing those trillion dollar secret ideas on model training.

well, because most people are not dickhead I guess?


Hold up. Look, this is all shades of grey but saying Chinese labs all release open weights stuff is kinda crazy thing to say.

Right now they are doing that because they are still trying to catch up to Anthropic, Google, and OpenAI.

The moment they have the special sauce, they will shut it down and you won't be able to run their stuff anymore outside of them. Why do I say that? We already have the evidence in the diffusion model arena. All the chinese labs were pumping out open weights models for image and video, the moment they got to SOTA, they stopped doing it. Less and less is being released.

Chinese companies aren't doing open weights models out of the goodness of their hearts, they are doing it because it help their entire industry catch up. Don't get it twisted, this is very much a US vs China battle here. China wants to win and I am not sure how they won't. Deepseek is the first major large model trained on Huawei chips. It won't be the last and I am betting that China will make up for lesser performance of those chips with more manufacturing and power generation.

I am very bullish on China winning the AI war here. But I also am not naive enough to think that the Chinese companies is doing open weights out of wanting to make the world a better place or the goodness of their hearts. It undercuts the american AI companies.


Now we get to the nub. American anti-Chinese rhetoric. Very good.


I made no such claims. Maybe you have something to share about why we need to have a negative view of free and open models based on publicly available frontier research.


I am personally okay helping them as long as they publish the models and dont keep them closed. And I dont trust the settings where providers say they wont train on it.


Because they give it away for free and offer APIs at very acceptable rates. Not that hard to figure out, Robin Hood stealing our data tax back comes to mind.


GitHub is free.


User publishes to github => Copilot trains with GitHub data => MS Sells copilot => User workes for Microsoft (in the sense of giving it's labour for MS to make money)

User publishes to github => Deepseek trains with GitHub data => Deepseek gives model away for free => User did not work for Deepseek (in the sense of giving it's labour for Deepseek to make money)


In the first case MS is giving part of Github itself away for free.


Exactly, it's intuitively different.


> I'm surprised that people here don't care at all about these models openly training on your data

You can use zero data retention and zero training providers for most open weights. See OpenRouter and OpenCode Go/Zen for examples.

This is actually one of the big selling points behind open weights - neither China nor the US get your data.


If they give me the resulting model in the end, they can train on my data all they want. Hell, I'll send them more of it.


If the data is opensource on github, then in my opinion it should be fair game.


IMO this is unfair for GPL or similarly licensed code.

Seems ok for MIT like licensed code though


It's totally fair to use GPL code, it just means all the models built by Anthropic, OpenAI, etc. using GPL-licensed source are themselves bound by the GPL. Plus, any works created downstream using those AI tools.

We're on the verge of a golden age of software as soon as someone finds a court with courage.


Ah, you have much more faith in the legal system than I do. It's nice to dream, though.


There's no difference. Either you need to follow the license or you don't. MIT has requirements still.


I think AI will create an open source dark age. Gradually, we'll see a lot less new good open source code. A gradual shift back to the proprietary world. Simmilar to the 1950-1990 period.


Why would giving more people software freedom and the ability to reverse engineer nonfree code result in a dark age?


The data is not open source. They have open weights but the source data is never open.


Things being public should not be enough. just because someone leaked your medical information to the public via a data breach should not make it fair game. There should be some rules.


I feel that's a false dichotomy. The code on github is freely available for people to read and learn from, leaked medical data isn't.


I feel that's a flase dichotomy. The code visible on github is freely available for anyone to read and learn from.


So would be your leaked medical record.

The point is not that this situation seems absurd. The point is that we need some point where we say whats ok or not.

And by ignoring licensing of public code already we moved it closer to the worse end of the spectrum


There are rules. I believe that search engine indexing follows these rules and that so called "training" is search engine indexing.

But a court may differ in the future.


My policy is that I don't allow agents to access all code. Some of it is shielded behind bind mounts. Maybe this is a pathetic, artisanal (or ego-driven), reaction of mine to the inevitable. I allow them to work on about 90% of the code (most codebases fully), with some code being considered too valuable to expose to the vendor. When data is involved, LLMs only get to see anonymized data.

This cute policy of mine won't affect anything though. The more we use the models, the more the models will replace this kind of work. Centralisation of power is inevitable; in Medival Europe, we used to have state & church ruling. In modern times but before the internet, it was probably state and banks. Maybe with ongoing digitization (bank offices disappearing) making banks less costly to operate; combined with with bank bailouts, maybe govenments will fully nationalize or at least banks will consolidate.

Then the AI companies will consolidate with the internet information and communication companies (Google/Meta for the US, and Alibaba/Tencent for China). Maybe we'll end up with a few de-facto governmental megacorps that rule in tandem and close cooperation with the formal government, who might handle mostly infra, utilities and the army. The megacorp would control narrative more and take more of a paternal role (educating and protecting the citizens, normally handled by formal governments).

Does this make sense?


AWS Bedrock has DeepSeek models running on their infrastructure. That should be enough to prevent training on user data (there's a markup compared to DeepSeek's pricing though).

And unfortunately AWS doesn't have prepaid billing, so you can't just give the internet access to your API key without getting FinDDoS'd.


The latest one available for serverless inference looks to be from 8 months (Deepseek v3.1), which is an eternity and far behind.


If anyone is looking for a solution in this space. Fire me an email, I have a partner whose focussed closely on that problem set!


At this point, that's kind of the reason I use open-weight models through the official providers when I can now.

There's some use cases I won't use a hosted model for, and will only do self hosted.

Otherwise, if they're going to keep releasing open-weight models, I'm going to keep giving them data.


I am fine with them training on my open source code (which is pretty bad but not the point, because they're providing the service for free). I will be super pissed if I pay for enterprise and they train on it though. I believe this is the opinion of majority programmers.


At least Moonshot (Kimi) says in the ToS that they train on your prompts when using their paid API.


What do you mean specifically? Data passed through OpenRouter? Or that they too indiscriminately ingest data all over the web? If the former, I assume it's just that anyone still using them just doesn't care where the data comes from. If the latter, well, it seems like every day there's some news on some new model from somewhere, and it takes dedication to complain every time. There's also the factor that I believe DeepSeek is more open with the model, while others keep it entirely proprietary, which feels fairer and (personally) is also less offensive.


As opposed to?

Do you really think OpenAI, Anthropic or any other entity in the same business respects your data?

The Chinese AI companies who release open weights actually deserve whatever input you give them. They are the reason why there is competition and not duopolies in the domain.


I think Google, and likely Anthropic, indeed do honor the settings chosen by the user. For Google in particular it'd be very surprising if they didn't. That's also why both do everything they can to trick users into allowing it.

OpenAI, I wouldn't be surprised if you were right.


You mean the same Anthropic, that wouldn't blink an eye at intentionally overcharging users hundreds of dollars just for having a HERMES.md file in a repo, would be above taking your data for... ethical reasons?


They also INTENTIONALLY gave people full refunds for that case.


unfortunately the history of these big tech companies has shown that they do not care about data privacy and are even willing to lie about it. but I guess its irrelevant, in practice you have to assume the worst anyway since there is no way to verify it


The models doesn’t get better by themselves. You’re naive.


I never see the output of the Claude or MS models without having to pay for the privilege. All the Chinese models are open weight and open source.


Are you implying chinese models training on my data is worse than OpenAI/Grok/Claude training on my data?


Two factors. First is anti-americanism (or at least anti-american-capitalism).

But the more important one is the social contract. Github came far before LLM era. The branding around it is being the storage of open source projects and many users want to it stay away from AI hype. You won't expect LLM providers to stay away from AI hype (duh) so it's less an issue for them.


From the EU side. I think we'll make a cost comparison between the US ( where it's leaders are doing weird shit against the EU and pro Russia) vs China ( who at least gives cheap models and doesn't actually tries to take over an entire European country).

US has too much influence atm. I'm ok with switching between "bullies".


thanks for the heads up on github


I am using DS4 via cortecs.ai. There is no training and it is GDPR-compliant. The flip side, it is expensive.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: