Skip to content

Author

Yingrui Ji

We have 3 of 12 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction

Oblique-view urban remote sensing imagery inevitably exhibits geometric projection displacements between building roofs and footprints, leading to significant distortions in spatial structure. Existing approaches either ignore these deformations or handle them implicitly within segmentation-based frameworks, where progress is dominated by general segmentation advances rather than improvements in geometric correction. In this work, we explicitly define roof-to-footprint offset vector (RFOV) extraction as an independent learning task that decouples geometric alignment from semantic segmentation. To support this task, we introduce the Oblique City dataset (ObliCity), the first large-scale benchmark that integrates high-resolution UAV imagery and globally distributed satellite data, covering diverse city morphologies and camera perspectives. Methodologically, we reformulate DragOSM into DragRoof, an ODE-based framework inspired by human annotation behavior. By simulating the continuous process of dragging roofs toward their footprints, DragRoof learns deterministic, geometry-consistent offset fields and adaptively determines convergence through an end token. Extensive experiments on ObliCity demonstrate that DragRoof achieves state-of-the-art RFOV extraction performance, requiring fewer inference steps while delivering superior directional and length accuracy. Our dataset and model establish a principled foundation for studying projection displacement correction in oblique remote sensing imagery. The source code and dataset will be avaliable at https://github.com/likaiucas/DragRoof.

Kai Li, Yupeng Deng, Li-Gao Deng et al. · 0 citations
Open access Sep 2026

TempFinRAG: Multimodal Temporal Retrieval-Augmented Generation for Point-in-Time Financial Question Answering

Financial question answering is often treated as document question answering, although financial evidence is both multimodal and time-dependent. Semantically equivalent facts expressed in narrative text, tables, page images, or Extensible Business Reporting Language (XBRL) should support consistent answers, whereas a disclosure may support a query only after becoming public. We formalise this combination as crossmodal evidence symmetry under a causal temporal boundary and introduce TempFinRAG, a multimodal temporal retrieval-augmented generation (RAG) framework for point-in-time financial question answering. Given a question, company, and as-of date, the framework enforces the information boundary defined by U.S. Securities and Exchange Commission (SEC) filing availability; aligns page layout, text, table structure, and XBRL facts; retrieves time-valid evidence; executes auditable financial calculations; and generates a cited answer. A verifier checks temporal validity, claim support, numerical consistency, and the need to abstain. We further introduce TempFinQA, a point-in-time evaluation protocol built from public filings and XBRL facts, and evaluate the framework on complementary evidence-grounded, numerical, conversational, and multi-table benchmarks. On TempFinQA, TempFinRAG improves answer accuracy from 66.7% to 78.9% over hybrid RAG while reducing temporal evidence leakage from 10.8% to 1.7% and hallucination from 17.3% to 7.9%. Reliable financial question answering therefore requires consistent treatment across evidence representations and deliberately asymmetric access across time.

Lanju Tao, Zheng-Ji Li, Ying-Rui Ji et al. · 0 citations
Preprint Aug 2026

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

MotionCraft is presented, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface to deliver temporally consistent, high-quality reconstructions under streaming constraints.

Rong Fu, Chun-Lei Meng, Yangcheng Zeng et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.