Comprehensively evaluating AI agents across interactive environments is difficult due to fragmented tasks, scaffolds, verifiers, and scoring rules. Unfortunately, existing efforts to unify these evaluations are limited in scale and domain, making costly reruns necessary and leaving available data incomparable. We introduce MESSIER, a unified corpus of 957,611 records spanning 30 benchmarks, 745 agents, 11,891 tasks, and 74,263 verifiers. MESSIER combines public evaluation results with new runs on six underrepresented professional and scientific benchmarks, standardizing their heterogeneous components into a common schema. Using this corpus, we show that frontier progress is uneven across benchmark groups, with function-calling evaluations largely saturated, programming improving fastest, and enterprise workflows remaining most challenging. Counterfactual rescoring further shows that strict all-pass scoring in multi-verifier tasks can alter agent rankings. Finally, we derive capability scores from our corpus that correlate with Epoch's Evaluation Capability Index rankings at Spearman \r{ho} = 0.84. The scores can also be estimated for subsets defined by domain, occupation, action space, or verifier type. In essence, MESSIER is a reusable resource for studying agent performance at scale, and a basis for designing better evaluations.
Stefan Krsteski, Charlotte Meyer, Guillaume Allegre et al.· 0 citations
Construction firms operate in knowledge-intensive and complex project environments where critical project knowledge is frequently fragmented across teams, documents, and digital systems. This fragmentation limits the systematic capture, structuring, and reuse of knowledge, which are fundamental processes for reducing knowledge loss, improving decision making, and enhancing project performance across organizational boundaries. Although effective knowledge mapping (KMp) is recognized as a valuable mechanism for organizing and disseminating such information, empirical evidence remains limited regarding how organizational, human, and technological factors interact to influence its effectiveness in construction contexts. This study addresses this gap by examining the interrelationships among these factors and their collective impact on KMp within construction firms. A quantitative, survey-based methodology was used to gather data from professionals within the construction industry. Structural equation modeling (SEM) using SPSS 23 and AMOS 24 was used to analyze the data and validate the measurement constructs and model fit. The study found that organizational frameworks, technological infrastructures, and human competencies significantly influence the effectiveness of KMp. Technological advancements were identified as a mediating factor, emphasizing the need for integration between digital tools and organizational culture to enhance knowledge-sharing processes. This study contributes to the knowledge management field by providing a systematic, data-driven perspective on the enablers of effective KMp. It extends the discourse on digital transformation and knowledge management, highlighting the importance of sociotechnical alignment between workforce capabilities and technological infrastructures. Future research should explore the role of artificial intelligence and machine learning in automating KMp processes. Practically, the findings provide construction firms with a structured strategy that integrates organizational alignment, workforce development, and technological investment to enhance KMp effectiveness and improve project decision making.
Safi Ullah, Xiaopeng Deng, Diana R. Anbar· Journal of construction engi...· 0 citations
A domain-specific legal artificial intelligence system for construction contract disputes via hybrid knowledge integration based on the retrieval-augmented generation (RAG) paradigm, integrating five core legal texts and 500 adjudication cases within a dual-engine architecture is proposed.
Ying Lu, Xin-Yun Shen, Yujing Wang et al.· Journal of construction engi...· 0 citations
TradeLens is introduced, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations, which reframe the evaluation of LLM-based trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-to-profit conversion.
Qiqi Duan, Changlun Li, Chen Wang et al.· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This review provides a comprehensive overview of covalent docking algorithms, systematically categorizing their approaches according to covalent bond formation, which primarily include tethered docking, biased docking, and dynamic covalent docking approaches.
Shi Li, Hongyan Du, Hui Zhang et al.· WIREs Computational Molecula...· 2 citations
This work introduces an advanced equivariant graph attention neural network specifically engineered to model long-range atomic electrostatic interactions with high precision, and improves the model's accuracy, generalization, and robustness in complex scenarios.
Qiaolin Gou, Qun Su, Ji-Ke Wang et al.· Journal of Chemical Informat...· 1 citation
Colorectal cancer (CRC) exhibits substantial molecular heterogeneity, necessitating the inference of subtype-specific driver genes and their interactions for drug-target discovery and precision oncology. Prior studies often fail to capture subtle, latent nonlinear regulatory mechanisms (dark causal relationships) driving tumor progression in specific subtypes. Here, we develop an explainable intelligence computational framework, Symbolic Trajectory-Embedded Dark Causal Interaction Inference (STE-DC2I), which combines symbolic trajectory embedding with historical prediction mechanisms to model nonmonotonic oscillatory dependencies between genes. Integrating single-cell transcriptomic and multiomics profiles from malignant epithelial subpopulations, STE-DC2I classifies CRC subtypes, reconstructs developmental trajectories, and uncovers interpretable subtype-specific driver genes with functional relevance. Unlike correlation-based and explicit causal approaches, STE-DC2I captures weak yet biologically critical regulatory signals, outperforming state-of-the-art methods in predicting subtype-specific CRC driver genes. Functional assays in CRC cell lines (in vitro) validated nine predicted driver genes, highlighting their therapeutic potential.This work systematically explores dark causal interactions between genes in CRC subtypes. STE-DC2I offers interpretable insights and a generalizable strategy for CRC drug-target discovery.
Meng Huang, Huijin Hu, Ming Li et al.· Journal of Chemical Informat...· 0 citations
This Account describes a computational and AI-driven ecosystem for structure-based covalent drug discovery and dives into a suite of cutting-edge, AI-driven computational methods, exploring the potential of deep learning in tasks such as molecular docking, covalent binding site prediction, and lead optimization.
Shi Li, Hongyan Du, Xujun Zhang et al.· Accounts of Chemical Researc...· 4 citations
There is a number $\psi$ such that for all $\varepsilon>0$ the probability that the value of the output neuron is in $[\psi - \varepsilon, \psi + \varepsilon]$ tends to 1 as $n$ tends to infinity.
A critical analysis of tools and trust mark frameworks intended to operationalize trustworthy AI (TAI), drawing on a comprehensive dataset from the OECD identifies significant asymmetries in ethical focus, lifecycle coverage, stakeholder targeting, and tool typology.
Michael Papademas, Xenia Ziouvelou, K. Karpouzis et al.· arXiv.org· 0 citations
This research highlights the heightened threats to data integrity and stakeholder trust in these evolving ecosystems through an intensive examination of the literature, initiating a pioneering discourse emphasizing fostering a foundation for developing secure and trustworthy Liquid AI environments.
M. Agbese, Niko Mäkitalo, Muhammad Waseem et al.· IoT· 6 citations· ⚡1
An observational study of 20,574 coding-agent sessions from 1,639 repositories across IDE and CLI workflows operationalizes misalignment as a breakdown made visible through developer pushback, and annotates each episode along four axes: form, cause, cost, and resolution.