SyncAI.news, a Varaisys broadcasting
NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset

Authors: Dhruv Nathawani, Shuoyang Ding US, Vitaly Lavrukhin US, Jane Polak Scowcroft US, Oleksii Kuchaiev US

NVIDIA continues releasing permissive datasets in support of the open ecosystem with 6 Million Multilingual Reasoning Dataset.

Continuing the success of the recent Nemotron Post-Training Dataset v1 release used in Llama Nemotron Super model, and our Llama Nemotron Post-Training Dataset release earlier this year, we’re excited to release the reasoning dataset translated into five target languages: French, Spanish, German, Italian, and Japanese.

The newly released NVIDIA Nemotron Nano 2 9B brings these capabilities to the edge with leading accuracy and efficiency with a hybrid Transformer–Mamba architecture and a configurable thinking budget—so you can dial accuracy, throughput, and cost to match your real‑world needs.

Model Highlights (TL;DR)

  • Model size: 9B parameters
  • Architecture: Hybrid Transformer–Mamba (Mamba‑2 + a small number of attention layers) for higher throughput at similar accuracy to Transformer‑only peers
  • Throughput: Up to 6× higher token generation than other leading models in its size class
  • Cost: Thinking budget lets you control how many “thinking” tokens are used—saving up to 60% lower reasoning costs
  • Target: Agents for customer service, support chatbots, analytics copilots, and edge/RTX deployments
  • Availability: The model weights are available on Hugging Face, you can try the endpoint on build.nvidia.com, and the model will be available as NVIDIA NIM for high throughput and low latency
  • License: nvidia-open-model-license

The release represents a significant step forward in our continued commitment to openness and transparency in model development and improvement. By releasing training data, in addition to the training tools and final model weights, NVIDIA supports continued improvement of open‑weight models.

What’s in the dataset and how we built it

Table 1: Ratio of discarded data (measured by bytes) by enforcing output format

Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News