Y

Yi Liu

Total Citations
110
h-index
4
Papers
3

Publications

#1 2608.01548v1 Aug 03, 2026

Emergence Invariance: From Symbolized Thought to Structural Control

Language-first intelligence is constrained by which distinctions enter its symbolic record, which mappings its language--interpreter--environment complex can execute, and which possibilities can be realized with finite resources. We formalize these limits through an effective interface $φ$, executable support $Π_{φ,H}$, a realization profile $M$, and resource-indexed families $\mathcal F_s(\mathcal J,M)$. They induce four nested capability levels: current reachability, budget realizability, asymptotic realizability, and structural capability, each with literal and risk-equivalent forms. On finite task spaces, universal bounded-loss dominance is characterized by inclusion of closed convexified risk envelopes; directed deficiencies and resource transformations quantify approximate simulation and realization burden. The decomposition $\mathcal R^*_{s,\mathcal J,M}=\mathcal R^*_{\mathcal J}+C_s(\mathcal J,M)$ separates structural limits from finite-resource compensation gaps. We develop dynamic control over these boundaries. Target risk identifies the lowest capability level that must change. Ockhamian control operates within inherited structural possibilities; Chattonian-containing control can produce endpoint structural novelty beyond the inherited decision class. Diminishing returns to within-structure computation and sustained structural-enrichment value yield a switching threshold, with switching costs creating hysteresis. Matched DeepSeek V4-Flash evidence shows: thinking raises pointer chasing from $0/16$ to $14/16$, exact observational and memory twins remain at their $0.5$ floors, restoring decisive memory raises performance from $0.5$ to $1.0$, and executable support determines whether reasoning can become effective. The framework organizes LLM emergence limits through structural capability, finite realizability, boundary control, and recursive meta-control.

Yi Liu
0 Citations
#2 2604.19149v1 Apr 21, 2026

How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning

Thinking LLMs produce reasoning traces before answering. Prior activation steering work mainly targets on shaping these traces. It remains less understood how answer tokens actually read and integrate the reasoning to produce reliable outcomes. Focusing on quantitative reasoning, we analyze the answer-to-reasoning attention and observe a benign self-reading pattern aligned with correctness, characterized by a forward drift of the reading focus along the reasoning trace and a persistent concentration on key semantic anchors, whereas incorrect solutions exhibit diffuse and irregular attention pattern. We interpret this as internal certainty during answer decoding, where the model commits to a viable solution branch and integrates key evidence. Following this, we propose a training-free steering method driven by Self-Reading Quality (SRQ) scores combining geometric metrics for process control with semantic metrics for content monitoring. SRQ selects data to build steering vectors that guide inference toward benign self-reading and away from uncertain and disorganized reading. Experiments show that our method yields consistent accuracy gains.

Hao-Yuan Chen Yi Liu Tao Zhang Chengfu Huo Wei Hu +1
0 Citations
#3 2604.11304v1 Apr 13, 2026

BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows

Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profession, we introduce BankerToolBench (BTB): an open-source benchmark of end-to-end analytical workflows routinely performed by junior investment bankers. To develop an ecologically valid benchmark grounded in representative work environments, we collaborated with 502 investment bankers from leading firms. BTB requires agents to execute senior banker requests by navigating data rooms, using industry tools (market data platform, SEC filings database), and generating multi-file deliverables--including Excel financial models, PowerPoint pitch decks, and PDF/Word reports. Completing a BTB task takes bankers up to 21 hours, underscoring the economic stakes of successfully delegating this work to AI. BTB enables automated evaluation of any LLM or agent, scoring deliverables against 100+ rubric criteria defined by veteran investment bankers to capture stakeholder utility. Testing 9 frontier models, we find that even the best-performing model (GPT-5.4) fails nearly half of the rubric criteria and bankers rate 0% of its outputs as client-ready. Our failure analysis reveals key obstacles (such as breakdowns in cross-artifact consistency) and improvement directions for agentic AI in high-stakes professional workflows.

Skyler Wang F. Guzm’an Elaine Lau Mark Ducker Ronak A. Chaudhary +22
2 Citations