Skip to content

Author

Hrant Davtyan

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

FailBench: How Reliable are VLMs at Judging Robot Task Success?

Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benchmarks offer limited evidence of cross-domain generalization. We introduce FailBench, a benchmark for robot failure detection comprising 2,197 manipulation attempts across 14 public sources (12 real-world, 2 simulated). In FailBench, 75% of failures occur naturally, and six real-world sources come from non-failure-detection datasets. Evaluating 13 VLM-based detectors, we find the best model achieves only 0.77 mean balanced accuracy. Notably, models fine-tuned for failure detection consistently underperform general-purpose VLMs and their own pretrained baselines. Performance depends heavily on required visual evidence: models approach saturation when outcomes depend on observable object motion, but degrade to near-chance (<0.60 balanced accuracy) on contact-intensive assembly tasks. Error analysis reveals a systematic bias toward predicting success under ambiguous evidence, which persists even with increased reasoning effort. Finally, we show that input-level intervention--spatially localizing and cropping outcome-relevant regions--improves the top detector by 2.4 percentage points without extra training.

Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan · 0 citations
#artificial intelligence Preprint Aug 2026

A Statistical Audit of Physical AI Benchmark Redundancy

A matrix of 51 models on 12 physical AI benchmarks, selected from a registry of 51 benchmarks and 152 models by reporting density is constructed, combining scores from model cards and benchmark papers with their own evaluation runs under each benchmark's official protocol.

Zaruhi Navasardyan, Hrant Davtyan · 0 citations
#natural language process... Preprint Aug 2026

Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning

This work reports on building and operating a retrieval-augmented legal assistant for Uzbek that must run in two regimes: a managed cloud service that maximizes answer quality within a per-token cost ceiling, and an on-premises deployment for clients whose legal data may not leave their infrastructure, restricting us to open-weight models on limited local hardware under latency constraints.

Tatul Danielyan, Mariam Avetisyan, Hrant Davtyan · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.