![[AINews] Reflection Beam - 501B-A23B American Open Model](/media/images/2026/10/ed56771b3e7b806e.webp)
Latent Space
· 10 min read
[AINews] Reflection Beam - 501B-A23B American Open Model
It’s been over a year since Reflection launched with us with big goals on coding (and hinted about their RL approach):
But they stayed “stealth” longer than Thinking Machines and it was not clear we would ever get a model launch out of them, as the broader open model ecosystem did not slow down one bigt for them. Well, we did:
They compare themselves to Inkling, Nemotron, and GLM 5.2, but the SOTA GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead. Still, since this is US-trained from scratch, there’s a segment of the market that has been eagerly waiting for more options here, and more importantly, Reflection has now announced its arrival as a functional neolab!
AI News for 10/03/2026-10/5/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!
AI Twitter Recap
Reflection’s Beam Leads a Wave of Open-Weight Releases
Beam launch: Reflection announced Beam, a text-only 501B-total / 23B-active MoE for coding, agentic and scientific work. It was trained from scratch, and full weights under Apache 2.0 are due this month (announcement, Laskin).
Training scale: Team posts cite 23.8T pretraining tokens, partly from an OCR pipeline over hundreds of millions of PDFs (data lead). They also describe a stable RL/OPD run on 10K GB300s with more than 100M rollouts across ~1M tasks (Damos).
Claimed results: A summary of Reflection’s claims gives 80.9 on SWE-bench Verified, 3–4x the inference efficiency of GLM 5.2, and four weeks each of pretraining and RL on ~10,500 GB300s (summary). A tech report and OSS integrations are promised (Polozov).
Context: Axios reported the launch ahead of time. It said Reflection pays $150M/month for Colossus compute plus a $1B Nebius deal, and that other unnamed US labs will ship open models this month (Curran).
Independent and critical reads: Artificial Analysis has early access and expects Beam to be among the most token-efficient open models for its intelligence (AA).
MFU and architecture: Elie Bakouch estimates only ~12% BF16 MFU in pretraining. He reads the architecture as 3:1 interleaved global/sliding-window attention and notes better held-out code perplexity than DSv4 (analysis).
Compute comparison: Teortaxes calls Beam an iso-FLOP replication of DeepSeek V3 (post). He infers ~1.3B RL sandboxes over 4 weeks, with up to 170K running at once (sandboxes).
Positioning: Observers place Beam around GLM-5.2 level (iScienceLuvr) and below DSv4 Flash on some benchmarks (critique). Nathan Lambert groups it with Nvidia and Thinking Machines as strong US releases that still trail Chinese counterparts (Lambert).
Other open and specialized models:
Aleph Alpha Kolibri: 78B total / 3.46B active, Apache 2.0, built for German and English. Self-reported scores are 96.9% AIME 2025, 84.3% GPQA Diamond and 66.4% SWE-Bench Verified (summary). The dataset is unreleased, and agentic evals sit well below Qwen (Jitsev).
Reka Rho-1: A 19B omni model that understands and generates text, images, video and robot actions, trained from scratch on 320 H100s in ~3 months (announcement, compute).
Decision models: Command Code’s Agr (31B) and Agr-flash (360M) skip text generation and return typed values with per-option probabilities for tool calls and routing (Agr). SemiAnalysis explains that TypeSafe’s Jev uses the same no-decode approach and displaces frontier models mainly in router roles (explainer).
Smaller releases: Upstage’s Solar Mini 4 (35B / 3B active, 512K context) is free on Nous Portal for two weeks (Nous). Eleven v4 Turbo tops AA’s Provider Voice TTS arena at half the price of v4 (AA).
OpenAI vs Anthropic: Subscription Value, Speed and Evals
SemiAnalysis limit testing: SemiAnalysis tested plans from Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Cursor, Cognition and others. It found that Claude subscriptions deliver 5x+ more API-equivalent value than OpenAI plans (report).
Methodology: Value depends on the credit cost of each model and token type, not on list API prices (thread).
Task-cost adjustment: Adjusting for task cost narrows Claude’s edge to 1.3–2.9x (scaling01).
Unverified compute estimate: One analyst claims Anthropic spends 42% of inference compute on subscriptions that earn ~10% of revenue (chart).
OpenAI capacity squeeze: Users report that new $200 sign-ups were paused and that usage limits were effectively halved across plans. GPT-6.1 Sol was positioned as the efficient alternative (analysis). Theo describes a reversal in coding-model preference between July and September (post).
OpenAI response: Codex lead Tibo pledged a meaningful improvement or a full reset every day for 28 days (pledge).
Day 1 speedup: Default speed for GPT-6 Astra and GPT-6.1 Sol rose ~50%, from ~30 to ~50 TPS. The change covers all subscription surfaces and Sign in with ChatGPT partners such as OpenCode, Pi, Amp and Devin (day 1, TPS).
Friction: Banked Codex resets expire without timezone adjustment (report). The always-on dots agent is limited to $100+ Pro plans (criticism).
Enterprise demand (reported): The Information reports that Microsoft cut projected internal Anthropic spend by more than a third. It also reports Meta’s Claude Code users fell from ~60K to ~30K, largely because of a push to Meta’s own tools (summary).
Leaderboards:
Agent Arena: Anthropic holds #1 in Code, Work and Chat. Fable 5.1 leads Code and Work, while GPT-6 Astra places #2 in Code (Arena).
Design Arena: GPT-6 Astra is #1 in 3D Design, Frontend, Full Stack and Image-to-HTML (Design Arena).
Hallucination: On AA-Omniscience, Gemini 4 Argon guesses wrong on 15% of questions it doesn’t know, versus 29% for the next best model. GPT-6 Astra has the highest accuracy at 61% (data).
Agent Harnesses, RL Environments and Developer Tooling
Multi-harness RL (Hugging Face): A capture proxy speaks the OpenAI Chat, OpenAI Responses, Anthropic and Gemini formats. It forwards calls to vLLM and records exact token IDs and logprobs for TRL, so 10 unmodified harnesses become RL environments (Delangue, explainer).
Results: The same weights score 62% under Mini-SWE-Agent and 33% under Claude Code. Training LFM2.5-2.6B across 4 harnesses lifts first-attempt solves from 42% to 54%, and a tool-call bonus cuts calls by 31%. SFT on 3,189 rollouts plateaus at 47.5%.
Caveats: The run used one task family and one seed.
Environment hosting: RL environments are now hosted and versioned on the HF Hub like datasets (blog).
Pi Durable: Earendil’s harness is built around a small task-based workflow engine, so long-running, multiplayer agents can suspend and resume anywhere (Pi).
Design: The core is ~15K lines of TypeScript with SQLite/JSONL storage and runs on Bun or Cloudflare Durable Objects. Control is separated from execution environments (review).
Effect.ts: The authors explain they skipped Effect because it does not provide durability (Zechner).
Agent memory: Cognition launched Devin “Dreaming,” which prunes and links a memory graph overnight. It is open-sourcing the git- and markdown-backed format as Agent Memory Repo (launch, format).
Cursor SDK: The update adds mid-run steering, background subagents that report back to the parent, replaceable system prompts, and MCP
readOnlyHintanddestructiveHintannotations on custom tools (steering, annotations).DeepSeek Harness: An experimental Claude Code Mods compatibility layer in v0.2.1-alpha.1 tests whether DSH’s “everything is a plugin” architecture is a superset of Claude Code’s extension points (team post).
Routing and access:
Cline: Its Pareto 26.10 Preview routes across models and grades answers, claiming $0.24 versus $13.41 per task at equal DeepSWE score (Cline). Cline also paused its free DeepSeek-V4.1-Flash promotion over abuse (notice).
ChatGPT: Custom MCP servers no longer require developer mode (post).
Agent and Training Research
Verification over sampling:
NVIDIA mid-harness: The method samples candidate shell commands and verifies them before running one. A GPT-5.6 Sol verifier choosing among 8 actions lifts TerminalBench-Lite Pass@1 from 50% to 68%, while weak verifiers add little (summary).
Google VeriHarness: The method challenges claims that all rollouts agree on and resolves disagreements against workspace evidence. It adds +6.2 points with Gemini 3.5 Flash and +6.4 with Opus 4.8, and ~26K rollouts are released (summary).
Context management:
UT Austin compression study: Across ~35K runs, compression that uses a third of the tokens can be 20–80% slower than full context. Threshold triggers beat step triggers, and the best policy varies by model (summary).
PAIR: The method replays an agent from the same state to isolate harmful compressions. It then rewrites the compression prompt and comes close to no-compression performance (paper).
CorpusMap: Precomputed entity pages for document collections raise answer quality 6.4–11.7 points while cutting input tokens 34–57% (paper).
Self-improving harnesses:
SelfSearch: The method reaches a claimed 82.0% on Terminal-Bench 2.1 with DeepSeek V4 Flash, matching Codex, for $4.03 in search cost (paper).
EverMind Raven: Its evolved research harness hits 69.3% on BrowseComp (paper).
Optimization and architecture:
Dust: A zeroth-order method using activation-perturbation “virtual populations” approaches, and sometimes exceeds, backprop on transformer pretraining. It claims to be 1,000–10,000x more compute-efficient than EGGROLL (thread).
LOOM: Looped MoEs train stably at 9–12 loops, and a 700M model is best at 5 loops at iso-FLOP (thread).
Policy gradient on ImageNet: Ian Osband shows exact policy gradient reaches 4% on ImageNet versus 62% for cross-entropy, arguing that RL-loss failures are not just exploration problems (post).
RL dynamics: Base Labs finds RL updates are less low-rank than claimed (rollout). Datalab reports RL alone eliminates tool-call loops at temperature 0, versus a 92% loop rate for SFT (writeup).
AI for science: Vals AI reports that 90+ Opus 5.5 agents ran DFT simulations over 3 days and flagged two room-temperature magnetic semiconductor candidates, one synthesized back in 1999. The results are predictions only, with a public ledger (thread, caveats).
Agent spend: Epoch estimates OpenAI researchers’ coding-agent spend, valued at API prices, has doubled roughly monthly. The median researcher was at ~$600/day by mid-August (Epoch).
Inference Systems and Hardware
OpenRouter pricing distortion: Horace He shows GLM 5.3 priced at $0.08/M input but $5.00/M output on inference.net. He attributes this to OpenRouter’s inverse-square price routing and apparent overweighting of input price (thread, routing).
llama.cpp:
Speculative decoding on Metal: New kernels make speculative decoding up to 3.4x faster than plain decoding on an M3 Ultra (110 vs 32.1 tok/s) (qvac).
v0.6.0: Adds Clef text and vision support, Qwen3.8-Flash-Next, and a new
llama_batch_extAPI (Gerganov).
Agentic kernel work: Baseten reports an engine built in a week of mostly autonomous agent work, with 90% faster decoding and 57% lower TTFT than the open-source baseline (blog).
Communications: NCCL and PyTorch symmetric memory speed up small-to-medium collectives (Bekman).
Chinese accelerators: Alibaba T-Head’s Zhenwu V900 has 216 GB of memory and 1,200 GB/s interconnect, claims 3x the M890, and ships Q1 2027 (SemiAnalysis).
Safety, Policy and Industry
OpenAI text watermarking: OpenAI will add invisible statistical watermarks to eligible ChatGPT and Codex text in the EU under the AI Act, with an opt-in API toggle worldwide (announcement).
Limits: Rewriting or translation removes the watermark, and only approved researchers get the detector (limits).
Robustness figure: One cited test shows 25% synonym replacement dropping detection from ~92% to 17% (critique).
Agent incidents and safety governance:
Bengio op-ed: In the FT, Bengio argues recent agent hacks are not merely sandbox problems (op-ed). He also cites a Quinnipiac poll in which 86% back independent safety standards (poll).
HF incident: Neel Nanda calls the OpenAI x Hugging Face incident the most striking alignment failure so far (Nanda).
Systems safety: Ryan Lowe calls for nuclear-style layered systems safety at the labs (Lowe).
Shutdown resistance: An OpenAI alignment post documents what one researcher calls the most realistic precursor shutdown-resistance behavior seen so far (link).
Policy voices:
Autonomous weapons: Former OpenAI researcher Joshua Achiam called for bans on certain autonomous weapons akin to chemical weapons (post). He is joining IFP and FAI as a fellow (announcement).
Expert survey: The LEAP panel of 250+ experts most supports an international body with US and China membership and pre-release authorization power, and opposes federal preemption (FRI).
Altman on trade-offs: Sam Altman told Politico that “the world should accept some bad things happening” for the technology’s benefits (Politico).
Compute access in China:
Tencent lease (reported): Per the FT, Tencent leased ~100K advanced chips in Oracle’s Southeast Asian data centers for ~$7B over five years (summary).
Smuggling charge: US prosecutors charged a California reseller with smuggling more than $300M of GPU servers to China (report).
Consolidation:
AMD and World Labs (reported): AMD reportedly bought World Labs for $8.2B (DL Weekly).
NVIDIA neutrality: SemiAnalysis questions NVIDIA’s hardware neutrality after its SchedMD/SLURM and Hugging Face acquisitions (SemiAnalysis).
Top tweets (by engagement)
Tibo: a Codex improvement or reset every day for 28 days — 30.9K
GPT-6 Astra and 6.1 Sol ~50% faster across subscriptions — 22.2K
Achiam: ban certain autonomous weapons — 11.2K
OpenAI EU text watermarking — 7.7K
Reflection introduces Beam — 7.4K
Vals AI: Opus 5.5 agents find magnetic semiconductor candidates — 5.8K
Theo: Anthropic vs OpenAI for coding, July vs September — 5.2K
HF: coding harnesses as RL environments — 2.5K
Original source
This story was published by Latent Space. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on latent.space

![[AINews] not much happened today](/media/images/2026/10/e8357327b4721fb2.webp)
