SyncAI.news, a Varaisys broadcasting
OpenAI o1-mini
ON

OpenAI News

· 1 min read

AI LabsOpenAI News

OpenAI o1-mini

We’re releasing OpenAI o1‑mini, a cost-efficient reasoning model. o1‑mini excels at STEM, especially math and coding—nearly matching the performance of OpenAI o1 on evaluation benchmarks such as AIME and Codeforces. We expect o1‑mini will be a faster, cost-effective model for applications that require reasoning without broad world knowledge.

Today, we are launching o1‑mini to tier 5 API users⁠(opens in a new window) at a cost that is 80% cheaper than OpenAI o1‑preview. ChatGPT Plus, Team, Enterprise, and Edu users can use o1‑mini as an alternative to o1‑preview, with higher rate limits and lower latency (see Model Speed⁠).

Optimized for STEM Reasoning

Large language models such as o1 are pre-trained on vast text datasets. While these high-capacity models have broad world knowledge, they can be expensive and slow for real-world applications. In contrast, o1‑mini is a smaller model optimized for STEM reasoning during pretraining. After training with the same high-compute reinforcement learning (RL) pipeline as o1, o1‑mini achieves comparable performance on many useful reasoning tasks, while being significantly more cost efficient.

When evaluated on benchmarks requiring intelligence and reasoning, o1‑mini performs well compared to o1‑preview and o1. However, o1‑mini performs worse on tasks requiring non-STEM factual knowledge (see Limitations⁠).

Math Performance vs Inference Cost

AIMEInference Cost (%)

Mathematics: In the high school AIME math competition, o1‑mini (70.0%) is competitive with o1 (74.4%)–while being significantly cheaper–and outperforms o1‑preview (44.6%). o1‑mini’s score (about 11/15 questions) places it in approximately the top 500 US high-school students.

Codeforces

90012581650Elo

HumanEval

90.2%92.4%92.4%Accuracy

Cybersecurity CTFs

20.0%43.0%28.7%Accuracy (Pass@12)

MMLU
0-shot CoT

92.3%90.8%85.2%88.7%

GPQA
Diamond, 0-shot CoT

77.3%73.3%60.0%53.6%

MATH-500
0-shot CoT

94.8%85.5%90.0%60.3%

DomainWin Rate vs GPT-4o (%)

Model Speed

Chat speed comparison

Original source

This story was published by OpenAI News. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on openai.com

Similar News