2607.25641v1 Jul 28, 2026 cs.CV

OmniPhys: 지식 그래프 기반 벤치마킹 및 텍스트-이미지 생성에서의 물리적 상식에 대한 집단 최적화

OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation

Zhizhen Liu
Zhizhen Liu
Citations: 45
h-index: 3
Hua-zeng Chen
Hua-zeng Chen
Citations: 1,507
h-index: 23
Yarong Lan
Yarong Lan
Citations: 34
h-index: 2
Mingchen Tu
Mingchen Tu
Citations: 1
h-index: 1
Yajing Xu
Yajing Xu
Citations: 202
h-index: 7
Jiaoyan Chen
Jiaoyan Chen
Citations: 235
h-index: 8
Yichi Zhang
Yichi Zhang
Citations: 554
h-index: 12
Jeff Z. Pan
Jeff Z. Pan
Citations: 291
h-index: 8
Wen Zhang
Wen Zhang
Citations: 227
h-index: 7

텍스트-이미지 모델은 놀라운 시각적 충실도를 보여주지만, 종종 기본적인 물리적 상식을 위반합니다. 기존의 벤치마크는 대개 거칠고 일반적인 설명을 사용하며, 특정 물리 법칙에 대한 이해도를 정확하게 평가하지 못합니다. 또한, 생성 과정의 높은 무작위성으로 인해 현재 프롬프트 최적화 방법은 '그라디언트 환상' 현상을 겪는데, 이는 최적화기가 체계적인 결함이 아닌 일시적인 시각적 오류에 의해 오도되는 현상입니다. 이러한 문제점을 해결하기 위해, 우리는 물리 지식 그래프를 기반으로 한 엄격한 벤치마크인 OmniPhys를 소개합니다. OmniPhys는 PhET 시뮬레이션을 표준 교육 과정과 연결하여 지식을 시나리오로 변환하는 파이프라인을 구축하고, 이원적인 검증 프로토콜을 통해 진단적 스트레스 테스트를 수행합니다. 또한, 우리는 물리적 일관성을 개별 최적화 문제로 취급하는 반복 프레임워크인 OmniPrompt를 제안합니다. OmniPrompt는 각 쿼리에 대해 K개의 무작위 이미지를 수집하여 쿼리별 피드백 버퍼를 생성하고, 메타 정책 업데이트 전에 B개 쿼리의 피드백을 병합하여 시드 및 쿼리 로컬 노이즈를 제거합니다. 12가지 대표적인 텍스트-이미지 모델에 대한 평가 결과, 보편적인 물리적 제약 조건이 드러났습니다. 실험 결과는 OmniPrompt가 다양한 모델 아키텍처에서 물리적 일관성을 크게 향상시키며, 진화된 메타 정책의 전송 가능성과 효과를 입증한다는 것을 보여줍니다. 코드와 데이터는 다음 링크에서 확인할 수 있습니다: https://github.com/zjukg/OmniPhys

Original Abstract

While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmarks often rely on coarse-grained descriptions, failing to diagnose the mastery of specific physical principles. Moreover, the high stochasticity of generative processes causes current prompt optimization methods to suffer from gradient hallucinations, where optimizers are misled by transient visual artifacts rather than systemic flaws. To address these challenges, we introduce OmniPhys, a rigorous benchmark of 1,551 samples grounded in a Physical Knowledge Graph. By aligning PhET simulations with standard curricula, OmniPhys operationalizes a knowledge-to-scenario pipeline that performs diagnostic stress tests via a dual-path verification protocol. We further propose OmniPrompt, an iterative framework that treats physical alignment as a discrete optimization problem. For each query, OmniPrompt aggregates K stochastic images into a per-query feedback buffer. Across training, it further merges feedback from batches of B queries before each meta-policy update, filtering seed and query-local noise. Evaluations across 12 representative text-to-image models reveal universal physical bottlenecks. Results demonstrate that OmniPrompt significantly enhances physical consistency across diverse backbones, proving the transferability and efficacy of our evolved meta-policies. The code and data are available at https://github.com/zjukg/OmniPhys

0 Citations
0 Influential
21 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!