洞察 AI 业界
最新进展
开源模型发布、各模型推理吞吐与价格、vLLM / SGLang 社区动态、学术与工业界资讯,全自动聚合。
Highlights
Intelligence
智能指数 · 越高越好模型性能榜单
| 模型 | 机构 | 智能指数 | 吞吐 tok/s | 首 token (s) | 输入/输出 /1M |
|---|---|---|---|---|---|
| Cogito v2.1 (Reasoning) | Deep Cogito | - | - | - | $1.25 / $1.25 |
| Mi:dm K 2.5 Pro Preview | Korea Telecom | - | - | - | $- / $- |
| EXAONE 4.5 33B (Non-reasoning) | LG AI Research | - | - | - | $- / $- |
| GPT-4o Realtime (Dec '24) | OpenAI | - | - | - | $- / $- |
| Claude Sonnet 5 (Adaptive Reasoning, High Effort) | Anthropic | - | 70 | 10.29 | $2.00 / $10.00 |
| GPT-3.5 Turbo (0613) | OpenAI | - | - | - | $- / $- |
| Claude Sonnet 5 (Adaptive Reasoning, Low Effort) | Anthropic | - | 60 | 1.83 | $2.00 / $10.00 |
| GPT-5.5 Pro (xhigh) | OpenAI | - | - | - | $- / $- |
| Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) | Anthropic | - | 65 | 2.33 | $2.00 / $10.00 |
| Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | - | 81 | 30.77 | $2.00 / $10.00 |
| GPT-4o mini Realtime (Dec '24) | OpenAI | - | - | - | $- / $- |
| Gemini 3 Deep Think | - | - | - | $- / $- | |
| GPT-5.4 Pro (xhigh) | OpenAI | - | - | - | $30.00 / $180.00 |
| Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Anthropic | 65.7 | 67 | 273.14 | $10.00 / $50.00 |
| Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Anthropic | 64.8 | 68 | 130.21 | $10.00 / $50.00 |
| Claude Opus 5 (Adaptive Reasoning, Max Effort) | Anthropic | 63.1 | 58 | 85.93 | $5.00 / $25.00 |
| Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | 62.5 | 55 | 33.78 | $5.00 / $25.00 |
| Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) | Anthropic | 62.5 | 56 | 31.83 | $10.00 / $50.00 |
| Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 62.1 | 70 | 100.45 | $10.00 / $50.00 |
| Muse Spark 1.3 (max) | Meta | 62.1 | - | - | $- / $- |
| Claude Opus 5 (Adaptive Reasoning, High Effort) | Anthropic | 61.5 | 55 | 19.49 | $5.00 / $25.00 |
| GPT-6 Astra (max) | OpenAI | 61.2 | - | - | $10.00 / $50.00 |
| GPT-6 Astra (xhigh) | OpenAI | 61.0 | - | - | $10.00 / $50.00 |
| Grok 4.6 (high) | SpaceXAI | 60.9 | 65 | 51.95 | $2.00 / $6.00 |
| GPT-5.6 Sol (max) | OpenAI | 60.9 | 81 | 98.87 | $4.00 / $20.00 |
| Muse Spark 1.3 (xhigh) | Meta | 60.8 | 150 | 42.49 | $1.25 / $4.25 |
| Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) | Anthropic | 60.5 | 56 | 13.29 | $10.00 / $50.00 |
| GPT-6 Astra (high) | OpenAI | 60.3 | - | - | $10.00 / $50.00 |
| Grok 4.6 (xhigh) | SpaceXAI | 60.0 | 64 | 52.96 | $2.00 / $6.00 |
| Kimi K3 (max) | Kimi | 59.7 | 39 | 4.89 | $3.00 / $15.00 |
开源模型发布
智能指数 47.3 · 输出吞吐 - tok/s
智能指数 56.7 · 输出吞吐 - tok/s
智能指数 61.2 · 输出吞吐 - tok/s
智能指数 60.3 · 输出吞吐 - tok/s
智能指数 55.3 · 输出吞吐 - tok/s
智能指数 59.2 · 输出吞吐 - tok/s
智能指数 61 · 输出吞吐 - tok/s
智能指数 60.8 · 输出吞吐 149.91 tok/s
智能指数 51.7 · 输出吞吐 313.48 tok/s
智能指数 56.6 · 输出吞吐 312.33 tok/s
智能指数 58.7 · 输出吞吐 305.52 tok/s
智能指数 62.1 · 输出吞吐 - tok/s
社区动态
# Release v5.16.1 This is a special release as we include GLM! (and a few small fixes) # GLM-5.3-Flash <img width="4239" height="2643" alt="image" src="https://github.com/user-attachments/assets/17bc9c29-758b-44c8-8230-42f945ded209" /> GLM-5.3-Flash, the first **natively multimodal model** in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Op
# Release v5.16.0 ## New Model additions ### Qwen4-Exp <img width="2241" height="693" alt="image" src="https://github.com/user-attachments/assets/c838b5ba-ffea-42da-baa9-3f66178e3671" /> Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE). GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual
# v0.28.0 ## Highlights This release features 584 commits from 270 contributors (76 new)! * **Kimi-K3 performance push**: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (#50484), fused FlashKDA decode and prefill kernels (#50654, #51311, #52458), SiTU activation support for MegaMoE (#50510), GEMM-RS for sequence parallelism (#52079), combined all-gathers with 1.5~3x kernel-level speedup (#51070), an adaptive speculative token budget deli
# Highlights *710 PRs from 212 contributors.* **New models in this release** (see the [cookbook](https://docs.sglang.io/cookbook) for all supported models): | Model | Type | PRs | Cookbook | |---|---|---|---| | Muse Glimmer | Autoregressive (Multimodal) | [#34262](https://github.com/sgl-project/sglang/pull/34262) | [link](https://docs.sglang.io/cookbook/autoregressive/Meta/MuseGlimmer) | | Intern-S2-Mobius | Autoregressive | [#33691](https://github.com/sgl-project/sglang/pull/33691)
# Patch release v5.15.1 This patch most notably solves a few issues with DFlash and MTP candidate generators, as well as an issue where images could sometimes not be processed on accelerator if using Lanczos filter. It contains the following commits: - Fix DFlash candidate token device mismatch with device_map="auto" (#47877) by @sywangyi and @Cyrilvallez - Align logit distributions for CandidateGenerators using sampling (#48007) by @Cyrilvallez - Fix MTP config when mlp_layer_types
This is a patch release on top of v0.27.0. - Support quantized DSpark Markov heads (#50424)
# vLLM v0.27.0 Release Notes ## Highlights This release features 561 commits from 242 contributors (64 new)! * **Kimi K3 support** with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option to shard the shared expert instead of replicating it (#50656). * **More new mode
# Release v5.15.0 ## New Model additions ### Meta Muse Glimmer Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases. Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it can be deployed to local setups for privacy-aware applications such as coding, document analysis, personal assistants, Claw- or Hermes-like setups. Muse Glimmer is a dense 30B parameter model consisting of: - 2B ViT-style encoder for
# Highlights *582 PRs from 194 contributors.* **Kimi K3 day-0 support**: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linear-attention layers interleaved with 24 MLA layers, and a MoonViT3d vision tower, shipping as a native MXFP4 checkpoint. SGLang serves it from day 0 with DCP, DSpark speculative decoding, chunked-prefill PP with TP decode, KDA-aware prefix caching, HiCache L2 over DCP, LoRA on the quantize
# vLLM v0.26.0 Release Notes ## Highlights This release features 411 commits from 212 contributors (61 new)! * **New Inkling model family** with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), MTP=1 speculative decoding (#48869), LoRA (#48884), and standard ModelOpt NVFP4 quantization (#48990). * **DeepSeek-V4 performance push** across vendors: a specialized routing kernel (2.94% E2E TPOT, #48660), `fused_topk
# Highlights *574 PRs from 169 contributors.* **DSpark: confidence-driven speculative decoding**: A new speculative algorithm. It drafts semi-autoregressively in blocks, then sizes each verify window from the draft's own confidence instead of a fixed draft length. Reaches **383.7 tok/s at accept length ~5** on DeepSeek-V4-Pro, TP8 on B300 (bs=1). Enable with `--speculative-algorithm DSPARK` and `SGLANG_RAGGED_VERIFY_MODE=compact`; tune the block with `--speculative-dspark-block-size` ([#30
# Patch release v5.14.1 This patch solves a few issues which appeared when integrating Inkling model, most notably an issue affecting models using EncoderDecoderCache during assisted generation. It also fixes an issue that could appear during prefill with StaticCache and sdpa without padding for Inkling which uses a position_bias. It contains the following commits: - Fix sdpa prefill with position_bias (#47359) by @Cyrilvallez - Fix assisted decoding for models with EncoderDecoder cache
学术 / 工业资讯
Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and dependency on a provider. We present compile by training, which turns a natural-language specification into a reusable neural function. At compile time, teacher models generate task-specific examples that are used to train a small adapter for a compact interpreter. The resulting function runs without the teachers and c
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3$\times$ longer yet no more accurate. We trace this to three deficiencies - incomplete error observation, limited search diversity, and unreliable selection - and propose ESPO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose clusters all training errors into structural patterns in one round; Propose genera
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operat
Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, th
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps improving for hundreds of steps and recovers most of full-data OPD's gain across task domains and model families. We explain this result through the states visited during training and t
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure an
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL reward, fusing the two signals within a single step: either as a \emph{weighted-additive combination} or a \emph{teacher-modulated rescaling} of the RL advantage. In this paper, we show that a simple two-stage scheme, OPD-then-RL, consistently outperforms pure O
MLLM-based embedding models remain limited in compositional retrieval, often failing to distinguish scenes containing the same concepts but different attribute-object bindings. Yet the same backbone can resolve such distinctions when used as a cross-attentive reranker, motivating us to distill its compositional judgments into the embedding model. We propose CORE, which synthesizes candidate lists spanning five compositional matching levels and introduces a Rank-KL objective that trains the embed
Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful repairs should also be minimal, reviewable, and faithful to the original implementation. We study over-editing, the tendency of a model to rewrite code beyond what is required to fix a bug. We construct an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch. Acros
Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternative decision trajectories implicitly available within multilingual decoding. This opacity becomes particularly problematic in low-resource dialect settings, where multiple linguistically valid realizations may differ in lexical authenticity, register, and structural stability. We propose reframing translation as a structured decision space explored by autonomous translation agents. Instead
Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large language model (LLM) brand recommendations, yet no standardized protocol exists for setting iteration counts, selecting stability metrics, or establishing reliability thresholds. Objective: We formalize the Dice Roll Method as a reusable protocol for repeated-query auditing of LLM brand recommendations, grounded in a generative model of temperature-scaled nucleus sampling. Methods: Total re
模型性能数据来自 Artificial Analysis,发布与社区动态来自 GitHub / arXiv,每日自动更新。