AI方向
| 排名 | 方向 | 前景 | 现在拥挤程度 | 个人做论文 |
|---|---|---|---|---|
| S1 | Agentic Post-training / Long-horizon RL | ★★★★★ | ★★★★ | ★★★★ |
| S2 | Adaptive Inference / Test-time Compute | ★★★★★ | ★★★ | ★★★★★ |
| S3 | Agent Memory / Lifelong Learning | ★★★★★ | ★★★ | ★★★★★ |
| S4 | Multimodal Reasoning + World Model | ★★★★★ | ★★★★ | ★★★ |
| S5 | Mechanistic Interpretability → Model Control | ★★★★☆ | ★★★ | ★★★★★ |
| S6 | 新生成范式:Diffusion Language Model | ★★★★☆ | ★★ | ★★★ |
| S7 | Agent Environment / Evaluation / Verifier | ★★★★☆ | ★★★ | ★★★★★ |
| S8 | AI for Science | ★★★★★ | ★★★ | ★★ |
| A | Agentic / Adaptive RAG | ★★★★ | ★★★★ | ★★★★ |
| A | LLM/VLM Systems / Serving | ★★★★★ | ★★★ | ★★★★ |
2026 年 8 月的论文趋势 + Google/Meta/Anthropic/OpenAI 的研究动向 + 仍然没解决的技术瓶颈
AI 的主线正在从「训练一个更强的静态模型」转向「训练一个会思考、会行动、会使用记忆、会自己分配计算资源、能长期完成任务的系统」。
所以真正值得关注的不是 SFT、LoRA、RAG、Agent 这些名词本身,而是背后的几个尚未解决的基础问题:
How to learn / how to reason / how to act / how to remember / how much to compute / how to verify。
1
下一像素预测
为什么一开始选择了这种范式 感觉有点奇怪 就是为什么一个一个的预测 语言 扩散语言有道理 但是还是说就是你按理说你现在说一个字 后面1w字以后的内容也应该是已经有固定概率的了 就跟平行世界一样 那我就想知道为什么这个范式是对的 有道理的 另外就是生成固定长度的什么时候会停止 为什么会停止
可以。我按你上一条问题真正关心的东西重新搜了一遍。如果只搜“next-pixel prediction”,会把问题搜窄;更合理的是把它放进 Visual Autoregressive Modeling + Visual Tokenization + UMM(Unified Multimodal Models)这条大谱系里。
现在文献大致已经把视觉 AR 分成 next-pixel → next-token → next-scale / patch / detail / continuous entity;而 UMM 又进一步问:理解和生成究竟应该共享一个 representation/objective,还是应该 decouple。(arXiv)
下面是我筛出来的 100 篇。我用 ★ 标出我认为你要真正理解这个问题最应该先看的论文。
一、最原始的问题:为什么要 Next-Pixel?1–10
这 10 篇看的是“把图像写成概率链”这件事是怎么一步步出现,又为什么后来大家想逃离 raw pixel。PixelRNN 真的是你刚才那个问题的直接起点。(arXiv)
-
★ Pixel Recurrent Neural Networks — 2016,真正的逐 pixel AR 起点。
-
Conditional Image Generation with PixelCNN Decoders — PixelCNN,把 RNN 换成 causal convolution。
-
PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications — 改进 pixel likelihood。
-
PixelSNAIL: An Improved Autoregressive Generative Model — causal conv + self-attention。
-
★ Image Transformer — Transformer 第一次非常直接地进入 autoregressive image modeling。
-
★ Generative Pretraining from Pixels (iGPT) — 2020,把 GPT 式 next-pixel 预训练用于视觉表示。
-
★ Neural Discrete Representation Learning (VQ-VAE) — 从 raw pixel 转向 discrete latent 的关键。
-
Generating Diverse High-Fidelity Images with VQ-VAE-2 — hierarchical discrete latent。
-
★ Taming Transformers for High-Resolution Image Synthesis (VQGAN) — “先 tokenize 图像,再像语言一样建模”的经典路线。
-
Autoregressive Image Generation Using Residual Quantization — RQ-VAE/RQ-Transformer,多层残差 token。
其中真正发生的范式转变是:
[
\text{next raw pixel}
\rightarrow
\text{next discrete visual token}.
]
因为 raw pixel sequence 太长,而且 pixel 并不是好的 semantic unit。(arXiv)
二、AR 到底能不能学“理解”?11–20
这一组对你尤其重要,因为它不是单纯生成,而是在问:
next-pixel / next-patch objective 会不会自己产生 semantic representation?
(arXiv)
-
★ MaskGIT: Masked Generative Image Transformer — 不再坚持严格 left-to-right,迭代并行预测。
-
★ MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis — generation + representation learning 早期统一尝试。
-
Rejuvenating image-GPT as Strong Visual Representation Learners — 为什么原始 iGPT 表征不够强,以及如何救回来。
-
★ Scalable Pre-training of Large Autoregressive Image Models (AIM) — Apple AIM;极重要,研究 visual AR scaling。
-
★ Denoising Autoregressive Representation Learning (DARL) — AR prediction + denoising,探索生成与表示学习结合。
-
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning — 把 AR representation learning 扩到视频。
-
An Empirical Study of Autoregressive Pre-training from Videos — 系统研究视觉 AR scaling 与 representation。
-
★ An Image is Worth 32 Tokens for Reconstruction and Generation (TiTok) — 极端压缩视觉 token。
-
MaskBit: Embedding-free Image Generation via Bit Tokens — visual token 从 codebook 进一步变成 bit。
-
★ Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction? — 和你刚才的问题最直接的一篇。
最后这篇尤其值得看。它在 2025 年重新问:
过去 next-pixel 不行,到底是 objective 错了,还是单纯 compute regime 不够?
作者专门做了 IsoFLOPs scaling,结论之一就是 raw next-pixel 的主要障碍可能更多是计算量,而不必然是范式从根上错误。(arXiv)
三、2024 之后:GPT 式视觉 AR 大复兴,21–30
(arXiv)
-
★ Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation (LlamaGen) — 最纯粹的“LLM next-token 搬到视觉”实验。
-
★ Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction (VAR) — 极关键。
-
STAR: Scale-wise Text-conditioned AutoRegressive Image Generation — next-scale 做 text-to-image。
-
★ Autoregressive Image Generation without Vector Quantization (MAR) — discrete token 不是必须的;用 diffusion loss 建模 continuous token。
-
★ Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens — continuous + random order。
-
Randomized Autoregressive Visual Generation (RAR) — 打破固定 raster ordering。
-
★ RandAR: Decoder-only Autoregressive Visual Generation in Random Orders — “生成顺序为什么非得固定?”
-
★ Next Patch Prediction for Autoregressive Visual Generation — 从 next-token → next-patch。
-
CART: Compositional Auto-Regressive Transformer for Image Generation — next-detail / compositional decomposition。
-
★ Next-X Prediction for Autoregressive Visual Generation (xAR) — 非常值得看:直接把“next 什么?”一般化。
xAR 的思想和你刚才的问题高度一致:所谓 (X) 可以是一个 patch、cell、subsample、scale,甚至 whole image。也就是说,autoregression 真正要求的不是“一个一个 pixel”,而只是选择一种 factorization。 (arXiv)
四、“到底应该预测什么?”31–40
这是目前我觉得最有研究味道的一组。
(arXiv)
-
★ HART: Efficient Visual Generation with Hybrid Autoregressive Transformer — discrete global structure + continuous residual。
-
★ Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis — bitwise AR + next-scale。
-
★ FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction — 挑战 VAR residual prediction。
-
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching — next-scale + flow。
-
★ DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation — diffusion trajectory 与 AR 融合。
-
D2C: Unlocking the Potential of Continuous Autoregressive Image Generation — discrete + continuous hybrid。
-
Frequency Autoregressive Image Generation with Continuous Tokens — 用 frequency structure 定义 autoregressive progression。
-
★ Parallelized Autoregressive Visual Generation — 直接攻击 AR sequential inference bottleneck。
-
Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots (Hi-MAR) — 先 global pivot,再 dense tokens。
-
★ DetailFlow: 1D Coarse-to-Fine Autoregressive Image Generation via Next-Detail Prediction — 把顺序定义成“detail 增长”。
这里已经能看到:
[
\text{next pixel}
]
根本不是 AR 的本体。
大家实际上正在搜索:
[
\boxed{
\text{What should be the autoregressive unit?}
}
]
Pixel?Token?Patch?Scale?Frequency?Detail?Continuous latent?Bit?Whole image?
这很可能才是你真正感兴趣的问题。(arXiv)
五、UMM 第一波:Understanding + Generation 怎么统一?41–50
UMM 的核心矛盾非常直接:
-
understanding 喜欢 semantic representation;
-
generation 又需要保留 pixel-level detail;
-
text 喜欢 discrete autoregression;
-
image generation 又长期更喜欢 diffusion/flow。
因此“统一”远不是把两个模型粘起来。(arXiv)
-
★ DreamLLM: Synergistic Multimodal Comprehension and Creation
-
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
-
★ Chameleon: Mixed-Modal Early-Fusion Foundation Models
-
ANOLE: An Open, Autoregressive, Native Large Multimodal Model for Interleaved Image-Text Generation
-
★ Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
-
★ Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
-
MonoFormer: One Transformer for Both Diffusion and Autoregression
-
★ VILA-U: A Unified Foundation Model Integrating Visual Understanding and Generation
-
★ Emu3: Next-Token Prediction is All You Need
-
★ Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
这十篇里有一个漂亮的争论:
Chameleon / Emu3:
尽量都 tokenize,统一 next-token。
Transfusion:
文本 next-token,图像 diffusion,但共享 Transformer。
Show-o:
AR + discrete diffusion。
Janus:
backbone 可以统一,但 understanding encoder 和 generation encoder 最好别硬统一。
这已经是四种完全不同的“统一哲学”。(arXiv)
六、UMM 第二波:2025 原生统一模型,51–60
(arXiv)
-
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
-
★ VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
-
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
-
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
-
★ MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
-
UGen: Unified Autoregressive Multimodal Model with Progressive Vocabulary Learning
-
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
-
★ BLIP3-o: A Family of Fully Open Unified Multimodal Models—Architecture, Training and Dataset
-
★ Emerging Properties in Unified Multimodal Pretraining (BAGEL)
-
★ MMaDA: Multimodal Large Diffusion Language Models
尤其 BAGEL 和 MMaDA 很值得放在一起看。
BAGEL 是 decoder-only native UMM,大规模 interleaved multimodal pretraining;MMaDA 则更激进,试图用 diffusion-style probabilistic formulation 同时处理 text reasoning、understanding 和 image generation。(arXiv)
七、我认为现在 UMM 最关键的问题:Visual Tokenizer,61–70
如果你以后想找创新点,这一组甚至比“大模型架构”更重要。
(arXiv)
-
★ TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
-
★ UniTok: A Unified Tokenizer for Visual Generation and Understanding
-
★ Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
-
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
-
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
-
UniCode²: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
-
★ UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
-
★ TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
-
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
-
Factorized Visual Tokenization and Generation
为什么重要?
因为 generation tokenizer 的目标往往是:
[
z \rightarrow \hat x \approx x
]
也就是尽量把 pixel detail 保存下来。
而 understanding encoder 更希望:
[
x\rightarrow z_{\text{semantic}}
]
满足:
两只姿势、光照不同的猫仍然应该非常接近。
所以:
[
\boxed{
\text{reconstruction optimal representation}
\neq
\text{semantic optimal representation}
}
]
这就是 UMM 最深的 representation conflict 之一。TokenFlow、UniTok、TUNA 等工作基本都在不同角度攻击它。(arXiv)
八、统一以后会不会“互相帮助”?71–80
另一个非常值得做研究的问题:
understanding + generation 放在一起,究竟产生 synergy,还是只是 multi-task interference?
(arXiv)
-
★ Unified Autoregressive Visual Generation and Understanding with Continuous Tokens (UniFluid)
-
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
-
Nexus-Gen: Unified Image Understanding, Generation, and Editing
-
★ Show-o2: Improved Native Unified Multimodal Models
-
HaploOmni: Unified Single Transformer for Multimodal Understanding and Generation
-
Skywork UniPic: Unified Autoregressive Modeling for Image Understanding, Generation and Editing
-
★ OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
-
EMMA: Efficient Multimodal Understanding, Generation and Editing
-
STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
-
UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning
UniFluid 很值得注意,因为它直接报告:
understanding 与 generation 确实存在 inherent trade-off,但合适的 loss balance 又能让两者互相促进。
这比单纯喊“统一更好”有研究价值得多。(arXiv)
九、另一派:为什么非得 Autoregressive?81–90
这直接对应你前面说的:
“既然后面一万个 token 都已经存在联合概率,为什么一定一个一个 sample?”
这些工作开始尝试用 diffusion / masked diffusion / flow 去直接建模更接近 joint 的对象。(arXiv)
-
★ Unified Multimodal Discrete Diffusion
-
★ Omni-Diffusion: Unified Multimodal Understanding and Generation
-
LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model
-
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
-
★ FlowInOne: Unifying Multimodal Generation as Image-in, Image-out Flow Matching
-
UniCanvas: A Diffusion-based Unified Model for Text-in-Image Generation
-
Uni-ViGU: Towards Unified Video Generation and Understanding
-
★ D-AR: Diffusion via Autoregressive Models
-
Distilled Decoding 1: One-step Sampling of Image Auto-Regressive Models with Flow Matching
-
MMaDA-Parallel: Multimodal Large Diffusion Models with Parallel Multimodal Interaction
其中 Omni-Diffusion 特别贴近你的“平行世界”想法:它明确不是单纯 left-to-right,而是用 masked discrete diffusion 去学习 multimodal discrete tokens 的 joint distribution。(arXiv)
十、2026 前沿:开始真正研究“统一是不是伪命题”,91–100
这是我认为你现在最值得追的一组。
(arXiv)
-
★ ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations
-
★ Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification (UniAR)
-
★ UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
-
★ Unified Sequential Modeling Activates Multimodal Understanding and Generation (NextFlow)
-
★ Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation (UniMRG)
-
★ Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal Models
-
Representation Forcing for Bottleneck-Free Unified Multimodal Models
-
HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
-
★ Vision as Unified Multimodal Generation
-
ILLUME-X: Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
91–97 已经明显进入一个新的阶段:
以前:
“把 understanding 和 generation 放进同一个模型。”
现在:
“它们为什么应该统一?统一哪一层?representation?tokenizer?backbone?objective?context?gradient?还是根本不应该全部统一?”
例如 UniAR 认为关键在 shared context-visual tokenizer;UniDDT 反而强调需要适当 decouple;UniMRG 又发现让模型生成 pixel/depth/segmentation 可以反过来改善 understanding。(arXiv)
我把这 100 篇压缩成一条发展逻辑
你不要把它们理解成 100 个孤立工作。
真正的历史是:
PixelRNN / PixelCNN
↓
为什么不直接预测 raw pixel?
↓
iGPT
↓
pixel AR 能不能 scale 出 representation?
↓
VQ-VAE / VQGAN
↓
那就把 image 变成 discrete token
↓
GPT式 next-token image generation
↓
LlamaGen / Emu3
↓
问题:视觉根本没有天然语言顺序
↓
RAR / RandAR
↓
问题:一个 token 一个 token 太慢,而且破坏2D结构
↓
VAR
↓
next-scale
↓
NPP / xAR / DetailFlow / Frequency AR
↓
到底应该 next 什么?
↓
continuous token: MAR / Fluid
↓
甚至 autoregression 本身是不是必须?
↓
MaskGIT / diffusion / flow / MMaDA / Omni-Diffusion
↓
与此同时:
理解模型使用 semantic feature
生成模型使用 reconstruction / VAE feature
↓
两种 representation 冲突
↓
Chameleon / Show-o / Janus / Transfusion
↓
UMM
↓
TokenFlow / UniTok / TUNA / UniAR
↓
真正的问题变成:
什么才是“统一的视觉状态”?
这里我认为有一个非常重要的研究空缺
读完这些论文之后,我现在反而觉得你上一条问题可以进一步变成一个很像顶会研究问题的 formulation:
[
\boxed{
\text{What is the proper factorization of visual probability?}
}
]
语言默认:
[
P(x)=\prod_tP(x_t|x_{<t})
]
因为语言的 temporal/sequential ordering 本身就相当自然。
但视觉不一定应该是:
[
pixel_1\rightarrow pixel_2\rightarrow\cdots
]
它可能应该是:
[
global\ semantics
\rightarrow
objects
\rightarrow
relations
\rightarrow
geometry
\rightarrow
appearance
\rightarrow
texture
]
也就是说:
不是 next-pixel,也不是 next-token,甚至不只是 next-scale,而是 next-semantic-state / next-visual-state。
这恰好把 VAR + UMM + visual tokenizer + JEPA/world model + generation/understanding synergy 几条现在看起来分散的研究线连接起来了。VAR 已经证明改变 factorization unit 会非常重要;UniAR/TUNA 等则说明 representation 本身是 UMM 的瓶颈。(arXiv)
如果你的目标是找一个能做论文的方向,我不会让你平均读这 100 篇。最值得先精读的是 20、14、21、22、24、25、27、30、43、45、46、49、50、58、59、61、63、68、91、92、93、95 这一串;它基本能把“next-pixel 为什么存在 → 为什么失败 → visual AR 怎么复兴 → UMM 为什么出现 → 现在真正卡在哪里”完整串起来。