Claudish is a bidirectional translator between English and a language called Claudish. It is powered by compiled neural programs with 0.6 billion parameters. The translator supports conversion from English to Claudish and from Claudish back to English.
programasweights.com
1 min
8/22/2026
DiffusionGemma is an experimental open-weight language model that generates text through discrete diffusion, refining 256-token blocks in parallel rather than producing tokens sequentially. The design is intended to avoid the decoding bottleneck of autoregressive language models. The model was created by fine-tuning the mixture-of-experts Gemma 4 model, which has 3.8 billion activated parameters and 25.2 billion total parameters. Its two-stage training process used fewer than 10% of the original autoregressive model’s total training-token budget. Supervised fine-tuning first taught bidirectional denoising; a second phase combined reinforcement learning and sampler distillation to improve generation quality and inference efficiency. Across its evaluation suite, DiffusionGemma generated about 20 tokens per forward pass and about 1,500 output tokens per second on a single Nvidia H100 GPU. The researchers say these results establish a new speed-capability trade-off frontier and exceed autoregressive models, including those using speculative decoding. DiffusionGemma retains Gemma 4’s thinking mode, multimodal-input, and long-context support. It can also still generate text autoregressively with minor performance degradation, suggesting potential hybrid diffusion-autoregressive decoding systems.
arxiv.org
2 min
8/20/2026
The paper "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" argues that large language models generate text by statistically predicting likely sequences of words. The publication gained significant attention following the firing of two authors, Timnit Gebru and Margaret Mitchell, by Google.
spectrum.ieee.org
10 min
7/6/2026
GPT-NL is a language model designed specifically for the Netherlands, emphasizing strong governance, transparency, and public values. It aims to integrate AI into various sectors such as the workplace, education, and public services while maintaining control over the technology.
tno.nl
3 min
6/16/2026
Large Language Models often utilize a construction known as negative parallelism, which establishes contrasts to reframe assumptions. This linguistic style is prevalent on social media platforms, particularly LinkedIn, and has generated significant discussion.
mail.cyberneticforests.com
10 min
5/31/2026
Mr. Chatterbox is a language model trained on over 28,000 Victorian-era British texts published between 1837 and 1899. The model can be run locally on personal computers and is based on a dataset provided by the British Library.
simonwillison.net
4 min
3/31/2026
Claudish is a bidirectional translator between English and a language called Claudish. It is powered by compiled neural programs with 0.6 billion parameters. The translator supports conversion from English to Claudish and from Claudish back to English.
programasweights.com
1 min
8/22/2026
The paper "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" argues that large language models generate text by statistically predicting likely sequences of words. The publication gained significant attention following the firing of two authors, Timnit Gebru and Margaret Mitchell, by Google.
spectrum.ieee.org
10 min
7/6/2026
Large Language Models often utilize a construction known as negative parallelism, which establishes contrasts to reframe assumptions. This linguistic style is prevalent on social media platforms, particularly LinkedIn, and has generated significant discussion.
mail.cyberneticforests.com
10 min
5/31/2026
Mr. Chatterbox is a language model trained on over 28,000 Victorian-era British texts published between 1837 and 1899. The model can be run locally on personal computers and is based on a dataset provided by the British Library.
simonwillison.net
4 min
3/31/2026
DiffusionGemma is an experimental open-weight language model that generates text through discrete diffusion, refining 256-token blocks in parallel rather than producing tokens sequentially. The design is intended to avoid the decoding bottleneck of autoregressive language models. The model was created by fine-tuning the mixture-of-experts Gemma 4 model, which has 3.8 billion activated parameters and 25.2 billion total parameters. Its two-stage training process used fewer than 10% of the original autoregressive model’s total training-token budget. Supervised fine-tuning first taught bidirectional denoising; a second phase combined reinforcement learning and sampler distillation to improve generation quality and inference efficiency. Across its evaluation suite, DiffusionGemma generated about 20 tokens per forward pass and about 1,500 output tokens per second on a single Nvidia H100 GPU. The researchers say these results establish a new speed-capability trade-off frontier and exceed autoregressive models, including those using speculative decoding. DiffusionGemma retains Gemma 4’s thinking mode, multimodal-input, and long-context support. It can also still generate text autoregressively with minor performance degradation, suggesting potential hybrid diffusion-autoregressive decoding systems.
arxiv.org
2 min
8/20/2026
GPT-NL is a language model designed specifically for the Netherlands, emphasizing strong governance, transparency, and public values. It aims to integrate AI into various sectors such as the workplace, education, and public services while maintaining control over the technology.
tno.nl
3 min
6/16/2026
Claudish is a bidirectional translator between English and a language called Claudish. It is powered by compiled neural programs with 0.6 billion parameters. The translator supports conversion from English to Claudish and from Claudish back to English.
programasweights.com
1 min
8/22/2026
GPT-NL is a language model designed specifically for the Netherlands, emphasizing strong governance, transparency, and public values. It aims to integrate AI into various sectors such as the workplace, education, and public services while maintaining control over the technology.
tno.nl
3 min
6/16/2026
Mr. Chatterbox is a language model trained on over 28,000 Victorian-era British texts published between 1837 and 1899. The model can be run locally on personal computers and is based on a dataset provided by the British Library.
simonwillison.net
4 min
3/31/2026
DiffusionGemma is an experimental open-weight language model that generates text through discrete diffusion, refining 256-token blocks in parallel rather than producing tokens sequentially. The design is intended to avoid the decoding bottleneck of autoregressive language models. The model was created by fine-tuning the mixture-of-experts Gemma 4 model, which has 3.8 billion activated parameters and 25.2 billion total parameters. Its two-stage training process used fewer than 10% of the original autoregressive model’s total training-token budget. Supervised fine-tuning first taught bidirectional denoising; a second phase combined reinforcement learning and sampler distillation to improve generation quality and inference efficiency. Across its evaluation suite, DiffusionGemma generated about 20 tokens per forward pass and about 1,500 output tokens per second on a single Nvidia H100 GPU. The researchers say these results establish a new speed-capability trade-off frontier and exceed autoregressive models, including those using speculative decoding. DiffusionGemma retains Gemma 4’s thinking mode, multimodal-input, and long-context support. It can also still generate text autoregressively with minor performance degradation, suggesting potential hybrid diffusion-autoregressive decoding systems.
arxiv.org
2 min
8/20/2026
Large Language Models often utilize a construction known as negative parallelism, which establishes contrasts to reframe assumptions. This linguistic style is prevalent on social media platforms, particularly LinkedIn, and has generated significant discussion.
mail.cyberneticforests.com
10 min
5/31/2026
The paper "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" argues that large language models generate text by statistically predicting likely sequences of words. The publication gained significant attention following the firing of two authors, Timnit Gebru and Margaret Mitchell, by Google.
spectrum.ieee.org
10 min
7/6/2026
No more articles to load