SyncAI.news, a Varaisys broadcasting
DeepMath: A lightweight math reasoning Agent with smolagents
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

DeepMath: A lightweight math reasoning Agent with smolagents

By Intel AI Software Group

DeepMath is an aligned math reasoning agent built on Qwen3-4B Thinking and fine-tuned with GRPO (Group Relative Policy Optimization). Instead of verbose text, the model emits tiny Python snippets for intermediate steps, runs them in a secure sandbox, and folds the results back into its reasoning, reducing errors and output length. The agent is implemented using the smolagents library.

We evaluate DeepMath on four math datasets: MATH500, AIME, HMMT, and HLE, and show that:

  • 🤖 The math agent alone reduces output lengths by up to 66%, while often improving accuracy.

  • ⚡ GRPO training improves the agent performance even further, in almost all benchmarks.

👉 Code and evaluation scripts: https://github.com/IntelLabs/DeepMath
👉 Model: https://huggingface.co/Intel/deepmath-v1

Why DeepMath?

Large language models (LLMs) have advanced reasoning capabilities, but mathematical problem-solving remains challenging; chain-of-thought traces can be lengthy and prone to arithmetic mistakes. Recent works[^1][^2] demonstrate that small models can reach strong performance, and other studies[^3] investigate tool use to improve reliability. What those papers generally do not emphasize is reducing trace verbosity or explicitly training models to prefer short, computation-oriented traces executed in a constrained, auditable environment.

We focused on two goals:

  1. Offload deterministic computation to a safe executor.

  2. Train models to prefer concise, computation-oriented traces over verbose text.

DeepMath tackles this by combining a small Python executor with a fine-tuned LLM, enabling concise, computation-driven reasoning. The model learns to generate short Python snippets, which are executed in a sandbox and reintegrated into the context. GRPO fine-tuning encourages this behavior by rewarding correctness and encouraging shorter outputs.

How It Works


Figure 2: Output example where python code is generated, evaluated and the answer is inserted into the trace and used for context.

Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News