Themata.AI
Themata.AI

Popular tags:

#developer-tools#ai-agents#llms#claude#ai-ethics#code-generation#ai-safety#openai#anthropic#discussion

AI is changing the world. Don't stay behind. Clear summaries, community insight, delivered without the noise. Subscribe to never miss a beat.

© 2026 Themata.AI • All Rights Reserved

Archive

|

Topics

|

Privacy

|

Cookies

|

Contact
🕒 Latest🔥 Top
WeekMonthYearAll Time

Filtering by tag:

cudaClear
Hot Chips 2026: CUDA Targets RISC-V
cudarisc-vgpu-computingnvidia
News

Hot Chips 2026: CUDA Targets RISC-V – By Chester Lam

Nvidia is extending CUDA support toward RISC-V CPUs, potentially allowing server-class RISC-V systems to host Nvidia GPU compute workloads. CUDA currently supports x86-64 and AArch64 CPUs. Nvidia requires RISC-V systems to use the RVA23 profile and comply with RISC-V server SoC and platform specifications, which cover baseline reliability, availability, serviceability and security features. Nvidia also requires hardware features beyond those specifications, including vector extensions, PCIe cache coherency and peer-to-peer PCIe communication between devices. PCIe coherency prevents GPU DMA transfers from missing modified CPU-cache data or leaving CPUs with stale cached data after GPU writes. Without it, CUDA software would need explicit cache invalidation handling. Peer-to-peer PCIe support allows device-to-device buffer transfers without routing data through CPU memory. ACPI support was another initial obstacle, but the UEFI Forum added RISC-V ACPI support in 2025, and the ratified RISC-V Boot and Runtime Services specification includes ACPI. Nvidia said it is partnering with SiFive, which plans to demonstrate a CUDA-capable system at Hot Chips; the cited example specifications indicate a high-core-count server processor. Nvidia also described NVLink Fusion, which lets partners integrate Nvidia NVLink IP with custom CPUs or accelerators, including potential RISC-V designs, subject to CUDA requirements and support for frameworks such as DOCA and NCCL.

chipsandcheese.com

🔥🔥🔥🔥🔥

5 min

8/24/2026

Nvidia is proposing a beast of a CPU system for Windows PCs

Nvidia is proposing a beast of a CPU system for Windows PCs. It has 128 GB of shared memory and comes with up to 6,144 state-of-the-art CUDA cores. CPU wise, the chip has 10 performance cores and 10 efficiency cores. The performance cores are based on the Cortex-X925. These Bạn mới làm quen với X ư? Đăng ký ngay để nhận dòng thời gian cá nhân hóa của riêng bạn!

twitter.com

🔥🔥🔥🔥🔥

1 min

6/6/2026

GitHub - Luce-Org/lucebox-hub: Lucebox optimization hub: hand-tuned LLM inference, built for specific consumer hardware.Tool

We got 207 tok/s with Qwen3.5-27B on an RTX 3090

Lucebox is an optimization hub for hand-tuned LLM inference, specifically designed for individual consumer hardware. It features kernels, speculative decoding, and quantization tailored for each target, with the first megakernel for hybrid DeltaNet/Attention LLMs achieving 1.87 tokens per joule on a 2020 GPU.

github.com

🔥🔥🔥🔥🔥

5 min

4/21/2026

BarraCUDA Open-source CUDA compiler targeting AMD GPUs

BarraCUDA is an open-source CUDA compiler designed for AMD GPUs, capable of compiling .cu files directly to GFX11 machine code and generating ELF .hsaco binaries. The compiler, written in 15,000 lines of C99, has no LLVM dependency and aims to support additional architectures in the future.

github.com

🔥🔥🔥🔥🔥

6 min

2/18/2026

Hot Chips 2026: CUDA Targets RISC-V – By Chester Lam

Nvidia is extending CUDA support toward RISC-V CPUs, potentially allowing server-class RISC-V systems to host Nvidia GPU compute workloads. CUDA currently supports x86-64 and AArch64 CPUs. Nvidia requires RISC-V systems to use the RVA23 profile and comply with RISC-V server SoC and platform specifications, which cover baseline reliability, availability, serviceability and security features. Nvidia also requires hardware features beyond those specifications, including vector extensions, PCIe cache coherency and peer-to-peer PCIe communication between devices. PCIe coherency prevents GPU DMA transfers from missing modified CPU-cache data or leaving CPUs with stale cached data after GPU writes. Without it, CUDA software would need explicit cache invalidation handling. Peer-to-peer PCIe support allows device-to-device buffer transfers without routing data through CPU memory. ACPI support was another initial obstacle, but the UEFI Forum added RISC-V ACPI support in 2025, and the ratified RISC-V Boot and Runtime Services specification includes ACPI. Nvidia said it is partnering with SiFive, which plans to demonstrate a CUDA-capable system at Hot Chips; the cited example specifications indicate a high-core-count server processor. Nvidia also described NVLink Fusion, which lets partners integrate Nvidia NVLink IP with custom CPUs or accelerators, including potential RISC-V designs, subject to CUDA requirements and support for frameworks such as DOCA and NCCL.

chipsandcheese.com

🔥🔥🔥🔥🔥

5 min

8/24/2026

We got 207 tok/s with Qwen3.5-27B on an RTX 3090

Lucebox is an optimization hub for hand-tuned LLM inference, specifically designed for individual consumer hardware. It features kernels, speculative decoding, and quantization tailored for each target, with the first megakernel for hybrid DeltaNet/Attention LLMs achieving 1.87 tokens per joule on a 2020 GPU.

github.com

🔥🔥🔥🔥🔥

5 min

4/21/2026

Nvidia is proposing a beast of a CPU system for Windows PCs

Nvidia is proposing a beast of a CPU system for Windows PCs. It has 128 GB of shared memory and comes with up to 6,144 state-of-the-art CUDA cores. CPU wise, the chip has 10 performance cores and 10 efficiency cores. The performance cores are based on the Cortex-X925. These Bạn mới làm quen với X ư? Đăng ký ngay để nhận dòng thời gian cá nhân hóa của riêng bạn!

twitter.com

🔥🔥🔥🔥🔥

1 min

6/6/2026

BarraCUDA Open-source CUDA compiler targeting AMD GPUs

BarraCUDA is an open-source CUDA compiler designed for AMD GPUs, capable of compiling .cu files directly to GFX11 machine code and generating ELF .hsaco binaries. The compiler, written in 15,000 lines of C99, has no LLVM dependency and aims to support additional architectures in the future.

github.com

🔥🔥🔥🔥🔥

6 min

2/18/2026

Hot Chips 2026: CUDA Targets RISC-V – By Chester Lam

Nvidia is extending CUDA support toward RISC-V CPUs, potentially allowing server-class RISC-V systems to host Nvidia GPU compute workloads. CUDA currently supports x86-64 and AArch64 CPUs. Nvidia requires RISC-V systems to use the RVA23 profile and comply with RISC-V server SoC and platform specifications, which cover baseline reliability, availability, serviceability and security features. Nvidia also requires hardware features beyond those specifications, including vector extensions, PCIe cache coherency and peer-to-peer PCIe communication between devices. PCIe coherency prevents GPU DMA transfers from missing modified CPU-cache data or leaving CPUs with stale cached data after GPU writes. Without it, CUDA software would need explicit cache invalidation handling. Peer-to-peer PCIe support allows device-to-device buffer transfers without routing data through CPU memory. ACPI support was another initial obstacle, but the UEFI Forum added RISC-V ACPI support in 2025, and the ratified RISC-V Boot and Runtime Services specification includes ACPI. Nvidia said it is partnering with SiFive, which plans to demonstrate a CUDA-capable system at Hot Chips; the cited example specifications indicate a high-core-count server processor. Nvidia also described NVLink Fusion, which lets partners integrate Nvidia NVLink IP with custom CPUs or accelerators, including potential RISC-V designs, subject to CUDA requirements and support for frameworks such as DOCA and NCCL.

chipsandcheese.com

🔥🔥🔥🔥🔥

5 min

8/24/2026

BarraCUDA Open-source CUDA compiler targeting AMD GPUs

BarraCUDA is an open-source CUDA compiler designed for AMD GPUs, capable of compiling .cu files directly to GFX11 machine code and generating ELF .hsaco binaries. The compiler, written in 15,000 lines of C99, has no LLVM dependency and aims to support additional architectures in the future.

github.com

🔥🔥🔥🔥🔥

6 min

2/18/2026

Nvidia is proposing a beast of a CPU system for Windows PCs

Nvidia is proposing a beast of a CPU system for Windows PCs. It has 128 GB of shared memory and comes with up to 6,144 state-of-the-art CUDA cores. CPU wise, the chip has 10 performance cores and 10 efficiency cores. The performance cores are based on the Cortex-X925. These Bạn mới làm quen với X ư? Đăng ký ngay để nhận dòng thời gian cá nhân hóa của riêng bạn!

twitter.com

🔥🔥🔥🔥🔥

1 min

6/6/2026

We got 207 tok/s with Qwen3.5-27B on an RTX 3090

Lucebox is an optimization hub for hand-tuned LLM inference, specifically designed for individual consumer hardware. It features kernels, speculative decoding, and quantization tailored for each target, with the first megakernel for hybrid DeltaNet/Attention LLMs achieving 1.87 tokens per joule on a 2020 GPU.

github.com

🔥🔥🔥🔥🔥

5 min

4/21/2026

No more articles to load