There would be nothing bad with LLMs provided that:
They weren’t based on colossal scale appropriation of people’s copyrighted work - are there any like that? The companies are actually laundering their IP addresses so that they can’t be blocked (see recent article on the LKML for an example). I don’t call that “fair use”.
They weren’t being powered by building new fossil-fuel generators in the middle of heatwaves that make it very clear that the climate emergency is here and now (the idea that LLMs are going to solve that problem for us is laughable).
I could go on. Nothing bad, so long as we ignore all the bad parts.
The closest to what you want is probably the Nvidia Nemotron series, which uses an open dataset. It used a sizable GPU cluster to train, but nothing on the scale of what OpenAI/Anthropic are guzzling. That, and Nvidia makes a point to advertise high utilization/efficiency.
There are smaller scale truly open LLMs like the Olmo series, but Nemotron is the most practical to use.
Then there are the Chinese “open weights” LLMs, which tend to be trained on more modest Huawei ASICs instead of GPUs, and probably with a good chunk of renewables. The dataset is closed, and who knows what is in there, but at least the training scale is much smaller
They are Apache licensed.
Another thing is that both these group pitch LLMs as modular tools to customize, not magic black boxes to rent.
There would be nothing bad with LLMs provided that:
They weren’t based on colossal scale appropriation of people’s copyrighted work - are there any like that? The companies are actually laundering their IP addresses so that they can’t be blocked (see recent article on the LKML for an example). I don’t call that “fair use”.
They weren’t being powered by building new fossil-fuel generators in the middle of heatwaves that make it very clear that the climate emergency is here and now (the idea that LLMs are going to solve that problem for us is laughable).
I could go on. Nothing bad, so long as we ignore all the bad parts.
The closest to what you want is probably the Nvidia Nemotron series, which uses an open dataset. It used a sizable GPU cluster to train, but nothing on the scale of what OpenAI/Anthropic are guzzling. That, and Nvidia makes a point to advertise high utilization/efficiency.
There are smaller scale truly open LLMs like the Olmo series, but Nemotron is the most practical to use.
Then there are the Chinese “open weights” LLMs, which tend to be trained on more modest Huawei ASICs instead of GPUs, and probably with a good chunk of renewables. The dataset is closed, and who knows what is in there, but at least the training scale is much smaller
They are Apache licensed.
Another thing is that both these group pitch LLMs as modular tools to customize, not magic black boxes to rent.