Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#discussion#llms#trending#claude#ai-ethics#code-generation#ai-safety#openai

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

Β© 2026 Themata.AI β€’ All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
πŸ•’ LatestπŸ”₯ Top

Filtering by tag:

foundation-modelsClear
Ornith-1.5: From Self-Scaffolding to Self-Improvement
foundation-modelsself-improvementreinforcement-learningai-research
Tool

Ornith-1.5: From Self-Scaffolding to Self-Improvement

Ornith introduced Ornith-1.5, a family of 397B-parameter mixture-of-experts (MoE), 35B MoE, and 9B dense models trained through a self-improvement loop. The system generates progressively harder tasks, creates task-specific scaffolds containing instructions, tools, decomposition, and orchestration, then produces reinforcement-learning rollouts. Rewards optimize task generation, scaffold construction, and solutions jointly using validity, frontier difficulty, novelty, solution quality, and resistance to reward hacking. The target task success rate is 0.2, favoring problems difficult enough to expose capability gaps while still yielding successful training trajectories. Ornith reports that its 397B model scored 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, compared with 85.0 and 59.0 for Claude Opus 4.8. It scored 86.0 on SWE-bench Verified, 92.8 on GPQA Diamond, and 80.0 on MCP-Atlas. The 35B MoE model activates 3B parameters per token and scored 68.5 on Terminal-Bench 2.1 using Claude Code and 79.0 on SWE-bench Verified. The 9B model scored 47.0 and 70.6 on those benchmarks, respectively; Ornith says its quantized Mobile version can run on iPhone and Android devices. Ornith reports that its benchmark results are averaged over five independent runs.

ornith.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

11 min

8/19/2026

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence.Research

Flux 3

FLUX 3 is a multimodal foundation model that learns from images, videos, and audio within a unified architecture. It aims to create a comprehensive representation of the world by understanding object interactions, movements, and corresponding sounds.

bfl.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

6 min

7/24/2026

Kimi K3, Qwen 3.8, and Anthropic's (potential) UnravellingNews

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

Moonshot Labs launched the Kimi K3 model, while Alibaba introduced the Qwen 3.8 model, both of which are reportedly close in performance to Anthropic's Fable 5. These models will have their weights released publicly in the coming weeks, posing a strategic challenge to leading model developers.

emergingtrajectories.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

5 min

7/20/2026

TabFM: A zero-shot foundation model for tabular data

TabFM is a new foundation model designed for tabular data, simplifying classification and regression workflows. It utilizes a "zero-shot" logic approach, similar to that of TimesFM, to enhance the handling of enterprise data infrastructure.

research.google

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

4 min

6/30/2026

Apertus – Open Foundation Model for Sovereign AI

Apertus Mini consists of 16 small language models that showcase distillation and quantization techniques. Developed by the Swiss AI Initiative, the models are fully open with documented training data, code, weights, and alignment principles, and are designed to comply with EU AI Act requirements by respecting opt-outs, removing PII, and preventing memorization.

apertvs.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

1 min

6/21/2026

Apple Foundation Models

Claude for Foundation Models is a Swift package that integrates Claude as a server-side language model within Apple's Foundation Models framework. It conforms to the LanguageModel protocol, allowing it to utilize the same LanguageModelSession API for functionality such as respond(to:), streaming, guided generation, and tool calling, with requests sent directly to the Claude API without Apple's involvement.

platform.claude.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

6 min

6/15/2026

Apple bets cheaper AI will woo small developers

Apple is offering lower AI infrastructure costs to attract newer developers, allowing those with fewer than 2 million first-time App Store downloads to use its Foundation Models in Private Cloud Compute without incurring cloud API costs. This initiative provides access to advanced AI capabilities while ensuring strong privacy protections.

techcrunch.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

2 min

6/8/2026

Apple reveals new AI architecture built around Google Gemini models

Apple announced a new AI architecture built on foundation models developed in collaboration with Google, utilizing technologies from the Gemini family. This architecture is designed to operate on-device and on servers through Apple's Private Cloud Compute infrastructure.

macrumors.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

1 min

6/8/2026

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

GLM-5V-Turbo is a foundation model designed for multimodal agents, enhancing their capabilities in language reasoning and perception across diverse contexts. The model aims to improve the performance of agents in real-world applications by integrating various modalities.

arxiv.org

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

2 min

5/5/2026

Google's 200M-parameter time-series foundation model with 16k context

TimesFM is a pretrained time-series foundation model developed by Google Research for time-series forecasting. The latest model version is TimesFM 2.5, with archived versions 1.0 and 2.0 available, and it can be accessed through the TimesFM Hugging Face Collection and is integrated into BigQuery as an official Google product.

github.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

2 min

3/31/2026

Ornith-1.5: From Self-Scaffolding to Self-Improvement

Ornith introduced Ornith-1.5, a family of 397B-parameter mixture-of-experts (MoE), 35B MoE, and 9B dense models trained through a self-improvement loop. The system generates progressively harder tasks, creates task-specific scaffolds containing instructions, tools, decomposition, and orchestration, then produces reinforcement-learning rollouts. Rewards optimize task generation, scaffold construction, and solutions jointly using validity, frontier difficulty, novelty, solution quality, and resistance to reward hacking. The target task success rate is 0.2, favoring problems difficult enough to expose capability gaps while still yielding successful training trajectories. Ornith reports that its 397B model scored 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, compared with 85.0 and 59.0 for Claude Opus 4.8. It scored 86.0 on SWE-bench Verified, 92.8 on GPQA Diamond, and 80.0 on MCP-Atlas. The 35B MoE model activates 3B parameters per token and scored 68.5 on Terminal-Bench 2.1 using Claude Code and 79.0 on SWE-bench Verified. The 9B model scored 47.0 and 70.6 on those benchmarks, respectively; Ornith says its quantized Mobile version can run on iPhone and Android devices. Ornith reports that its benchmark results are averaged over five independent runs.

ornith.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

11 min

8/19/2026

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

Moonshot Labs launched the Kimi K3 model, while Alibaba introduced the Qwen 3.8 model, both of which are reportedly close in performance to Anthropic's Fable 5. These models will have their weights released publicly in the coming weeks, posing a strategic challenge to leading model developers.

emergingtrajectories.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

5 min

7/20/2026

Apertus – Open Foundation Model for Sovereign AI

Apertus Mini consists of 16 small language models that showcase distillation and quantization techniques. Developed by the Swiss AI Initiative, the models are fully open with documented training data, code, weights, and alignment principles, and are designed to comply with EU AI Act requirements by respecting opt-outs, removing PII, and preventing memorization.

apertvs.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

1 min

6/21/2026

Apple bets cheaper AI will woo small developers

Apple is offering lower AI infrastructure costs to attract newer developers, allowing those with fewer than 2 million first-time App Store downloads to use its Foundation Models in Private Cloud Compute without incurring cloud API costs. This initiative provides access to advanced AI capabilities while ensuring strong privacy protections.

techcrunch.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

2 min

6/8/2026

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

GLM-5V-Turbo is a foundation model designed for multimodal agents, enhancing their capabilities in language reasoning and perception across diverse contexts. The model aims to improve the performance of agents in real-world applications by integrating various modalities.

arxiv.org

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

2 min

5/5/2026

Flux 3

FLUX 3 is a multimodal foundation model that learns from images, videos, and audio within a unified architecture. It aims to create a comprehensive representation of the world by understanding object interactions, movements, and corresponding sounds.

bfl.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

6 min

7/24/2026

TabFM: A zero-shot foundation model for tabular data

TabFM is a new foundation model designed for tabular data, simplifying classification and regression workflows. It utilizes a "zero-shot" logic approach, similar to that of TimesFM, to enhance the handling of enterprise data infrastructure.

research.google

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

4 min

6/30/2026

Apple Foundation Models

Claude for Foundation Models is a Swift package that integrates Claude as a server-side language model within Apple's Foundation Models framework. It conforms to the LanguageModel protocol, allowing it to utilize the same LanguageModelSession API for functionality such as respond(to:), streaming, guided generation, and tool calling, with requests sent directly to the Claude API without Apple's involvement.

platform.claude.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

6 min

6/15/2026

Apple reveals new AI architecture built around Google Gemini models

Apple announced a new AI architecture built on foundation models developed in collaboration with Google, utilizing technologies from the Gemini family. This architecture is designed to operate on-device and on servers through Apple's Private Cloud Compute infrastructure.

macrumors.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

1 min

6/8/2026

Google's 200M-parameter time-series foundation model with 16k context

TimesFM is a pretrained time-series foundation model developed by Google Research for time-series forecasting. The latest model version is TimesFM 2.5, with archived versions 1.0 and 2.0 available, and it can be accessed through the TimesFM Hugging Face Collection and is integrated into BigQuery as an official Google product.

github.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

2 min

3/31/2026

Ornith-1.5: From Self-Scaffolding to Self-Improvement

Ornith introduced Ornith-1.5, a family of 397B-parameter mixture-of-experts (MoE), 35B MoE, and 9B dense models trained through a self-improvement loop. The system generates progressively harder tasks, creates task-specific scaffolds containing instructions, tools, decomposition, and orchestration, then produces reinforcement-learning rollouts. Rewards optimize task generation, scaffold construction, and solutions jointly using validity, frontier difficulty, novelty, solution quality, and resistance to reward hacking. The target task success rate is 0.2, favoring problems difficult enough to expose capability gaps while still yielding successful training trajectories. Ornith reports that its 397B model scored 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, compared with 85.0 and 59.0 for Claude Opus 4.8. It scored 86.0 on SWE-bench Verified, 92.8 on GPQA Diamond, and 80.0 on MCP-Atlas. The 35B MoE model activates 3B parameters per token and scored 68.5 on Terminal-Bench 2.1 using Claude Code and 79.0 on SWE-bench Verified. The 9B model scored 47.0 and 70.6 on those benchmarks, respectively; Ornith says its quantized Mobile version can run on iPhone and Android devices. Ornith reports that its benchmark results are averaged over five independent runs.

ornith.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

11 min

8/19/2026

TabFM: A zero-shot foundation model for tabular data

TabFM is a new foundation model designed for tabular data, simplifying classification and regression workflows. It utilizes a "zero-shot" logic approach, similar to that of TimesFM, to enhance the handling of enterprise data infrastructure.

research.google

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

4 min

6/30/2026

Apple bets cheaper AI will woo small developers

Apple is offering lower AI infrastructure costs to attract newer developers, allowing those with fewer than 2 million first-time App Store downloads to use its Foundation Models in Private Cloud Compute without incurring cloud API costs. This initiative provides access to advanced AI capabilities while ensuring strong privacy protections.

techcrunch.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

2 min

6/8/2026

Google's 200M-parameter time-series foundation model with 16k context

TimesFM is a pretrained time-series foundation model developed by Google Research for time-series forecasting. The latest model version is TimesFM 2.5, with archived versions 1.0 and 2.0 available, and it can be accessed through the TimesFM Hugging Face Collection and is integrated into BigQuery as an official Google product.

github.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

2 min

3/31/2026

Flux 3

FLUX 3 is a multimodal foundation model that learns from images, videos, and audio within a unified architecture. It aims to create a comprehensive representation of the world by understanding object interactions, movements, and corresponding sounds.

bfl.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

6 min

7/24/2026

Apertus – Open Foundation Model for Sovereign AI

Apertus Mini consists of 16 small language models that showcase distillation and quantization techniques. Developed by the Swiss AI Initiative, the models are fully open with documented training data, code, weights, and alignment principles, and are designed to comply with EU AI Act requirements by respecting opt-outs, removing PII, and preventing memorization.

apertvs.ai

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

1 min

6/21/2026

Apple reveals new AI architecture built around Google Gemini models

Apple announced a new AI architecture built on foundation models developed in collaboration with Google, utilizing technologies from the Gemini family. This architecture is designed to operate on-device and on servers through Apple's Private Cloud Compute infrastructure.

macrumors.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

1 min

6/8/2026

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

Moonshot Labs launched the Kimi K3 model, while Alibaba introduced the Qwen 3.8 model, both of which are reportedly close in performance to Anthropic's Fable 5. These models will have their weights released publicly in the coming weeks, posing a strategic challenge to leading model developers.

emergingtrajectories.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

5 min

7/20/2026

Apple Foundation Models

Claude for Foundation Models is a Swift package that integrates Claude as a server-side language model within Apple's Foundation Models framework. It conforms to the LanguageModel protocol, allowing it to utilize the same LanguageModelSession API for functionality such as respond(to:), streaming, guided generation, and tool calling, with requests sent directly to the Claude API without Apple's involvement.

platform.claude.com

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

6 min

6/15/2026

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

GLM-5V-Turbo is a foundation model designed for multimodal agents, enhancing their capabilities in language reasoning and perception across diverse contexts. The model aims to improve the performance of agents in real-world applications by integrating various modalities.

arxiv.org

πŸ”₯πŸ”₯πŸ”₯πŸ”₯πŸ”₯

2 min

5/5/2026