![[Paper Review] NVIDIA Nemotron 3: Efficient and Open Intelligence](https://developer-blogs.nvidia.com/wp-content/uploads/2025/12/image3-8-png.webp)
[Paper Review] NVIDIA Nemotron 3: Efficient and Open Intelligence
Paper NVIDIA Nemotron 3: Pushing the “Accuracy/Throughput” Frontier with a Hybrid Mamba–Transformer MoENemotron 3 combines an …
4 min
Natural Language Processing
Machine Learning
Mixture of Experts
![[Paper Review] Radial Attention: O(n log n) Sparse Attention with Energy Decay for Long Video Generation](https://moonlight-paper-snapshot.s3.ap-northeast-2.amazonaws.com/arxiv/radial-attention-onlog-n-sparse-attention-with-energy-decay-for-long-video-generation-2.png)
![[Paper Review] Pretraining Large Language Models with NVFP4](https://developer-blogs.nvidia.com/wp-content/uploads/2025/08/Optimizing-LLM-Training-png.webp)