Artificial intelligence (AI) has fundamentally transformed online discourse, serving simultaneously as a source of and a potential solution to threats to democracy. This article conceptually examines AI’s multifaceted role in digital public spheres, beginning with an analysis of how AI shapes democratic processes online, ranging from information curation and political participation to dialogue facilitation and the construction of epistemic infrastructures. We identify several key risks: AI amplifies harmful content through automated content generation, coordinated manipulation, and the algorithmic mainstreaming of borderline content that often evades traditional detection systems. Within this broader democratic context, we position counter speech as a critical intervention strategy and examine both reactive measures (e.g., automated detection and content moderation) and proactive approaches (e.g., prebunking, algorithmic downranking, friction design, and AI-mediated dialogue). We argue that effective counter speech increasingly relies on human–AI collaboration, where AI supports human counter speakers with factual resources, emotional scaffolding, and scalability, while human actors preserve the authenticity and agency that make counter speech normatively meaningful. However, realizing the potential of such collaboration requires moving beyond technological fixes toward democratically legitimated governance structures. Drawing on examples from Wikipedia and decentralized platforms, we demonstrate how transparent and participatory institutional arrangements can foster more resilient discourse environments. We conclude that societies can harness AI’s democratic potential while mitigating its risks only by proactively establishing legitimate deliberative processes grounded in broadly shared discourse norms.
Diana Rieger, Mario Haim· Cogitatio (Cogitatio)· 0 citations
Modern industrial systems require increasing efficiency and flexibility to remain competitive. To achieve this, decision support systems are essential for assisting humans in steering manufacturing, helping operators meet productivity and quality goals while minimizing the environmental impact. Digital Twins (DTs) are a key enabling technology for these systems, as they can monitor the process and predict key performance indicators (KPIs) such as energy consumption and product quality across different scenarios. Conventionally, physics-based (PB) models are regarded as the gold standard for building DTs thanks to their interpretability and reliability. However, as industrial systems grow in complexity, modeling from first principles becomes prohibitive. Data-driven regression models offer a practical alternative by leveraging Machine Learning (ML) techniques to model the system directly from data. ML techniques can automatically capture general non-linear input-output relationships, which makes them significantly more flexible. Significant efforts have been made to digitalize factories within the Industry 4.0 paradigm, a prerequisite for training ML-based models. However, even with digitalization, a paradox arises: truly informative data remains scarce. Industries are typically risk-averse, limiting the exploration of new operating regimes to protect safety and maintain a stable production. As a result, historical data is typically confined to a small range and number of conditions, while the models must operate reliably in unseen scenarios to support human decision-making. Ironically, these are the scenarios where standard ML-based regression models lack trustworthiness, often lacking the robustness and physical consistency required to operate safely. This thesis therefore investigates how data-driven regression models can be made more trustworthy by design, by integrating prior knowledge and automatically selecting robust models, with a focus on building DTs for industrial machines. This research aligns with the EU Ethics Guidelines for Trustworthy Artificial Intelligence (AI) [55]. While these guidelines outline seven key requirements for trustworthiness, here the scope is limited to three pillars: human agency and oversight, technical robustness and safety, and transparency. Addressing these three pillars, the first part of the thesis focuses on integrating prior physical knowledge and human rationale into the modeling process. It starts by drawing theoretical connections between diverse practices in physics-informed ML, making it easier for practitioners to choose between methods. Specifically, it demonstrates that adding physical equations as input features is a special case of architectural inductive biases, and that data augmentation, with synthetic data generated prior to training, is a special case of soft constraints, the two becoming equivalent when the prior knowledge can be written explicitly as target values. It then proposes a framework designed to embed human decision-making rationale directly into the architecture and training process of models used for continuous manufacturing. Specifically, the system's time-series behavior is reduced to a static mapping. This preserves partial interpretability and aligns with operators' mental models, while enabling plug-and-play deployment over existing legacy control systems. To ensure physical consistency, the framework induces known monotonic relationships by regularizing the model's Jacobian matrix. It also emphasizes sensitivity to decision variables, a critical aspect often overlooked in conventional approaches. The framework is validated both on historical data and during a seven-month real-time deployment at a large-scale wooden fiberboard manufacturer. There, DTs were designed for five physically distinct process stages using the same approach, demonstrating its generality. The results show that the proposed framework enhances both model responsiveness and physical consistency compared to standard training, which often yields models with negligible sensitivity and limited practical utility. Extending the focus on technical robustness, the second part of the thesis complements the first by improving the robustness of data-driven regression models against distribution shifts by design. It addresses this challenge through a novel data splitting strategy named Leave-Boundary-Out (LBO). First, theoretical foundations are established, showing that out-of-distribution (OOD) regions can be exploited for more effective model selection. Because the sensitivity of the validation loss to hyperparameters is often amplified in these regions, sample-size requirements are reduced. The LBO algorithm then leverages this insight to select models with superior extrapolation performance. Validated on multiple synthetic and real-world benchmarks from diverse engineering domains, the results using LBO show that models tuned for extrapolation consistently outperform standard approaches in OOD scenarios. Importantly, for the Polynomial Lasso pipeline, models simpler than those tuned using standard approaches tend to extrapolate better under both the L0- and L1-norm complexity measures, with LBO acting as a regularization mechanism that favors OOD robustness over purely fitting in-distribution data; for SVR the trend was ambiguous. Additionally, LBO helps localized methods such as Radial Basis Function (RBF) kernel-based regressors learn more generalizable trends, which could prove valuable for optimization frameworks that largely rely on such models. Overall, this thesis provides theoretical, methodological, and experimental foundations demonstrating that trustworthy data-driven regression modeling requires tailored approaches based on use-case needs. Furthermore, it offers validated directions for achieving this in industrial and engineering contexts, such as DT development for process and product optimization.
Machine intelligence has conquered the symbolic world but stalled at the physical one. The stall is structural: physical AI faces a cold-start deadlock -- no intelligence without data, no data without deployed intelligence. Our thesis: the deadlock is real but unevenly distributed, and the exception has a name: the artificial physical world. Buildings, industrial facilities, and infrastructure are intentionally constituted and documented: designed artifacts ship with readable archives that precede and constitute their instances; here, norms are promulgated before instances, not averaged from them. Four contributions. (i) From a four-world ontology we derive a legitimacy criterion for constitutive prior frameworks: prior extraction is legitimate if and only if the object domain is intentionally constituted and has left a readable archive; the criterion is testable through direction of fit -- deviation from a constitutive norm is a violation in the world, not a revision of the model. (ii) We establish a layering lower bound: any such framework has at least four layers -- syntax, concept, knowledge, instance -- because four construction goals pair into mutually incompatible carriers. (iii) We register deployment claims across five industrial domains and a 32-class failure-mode vocabulary. (iv) We stake the framework on five falsifiable predictions, the central one checkable on the public engineering record: if it fails, the framework fails. Semi-formal arguments back these claims (Appendix A): a Gold-type boundary on rule coverage in archiveless worlds, a decidability result for failure reduction over closed concept layers, and a boundary theorem for certificate-anchored calculi. Large language models find an honored place here -- as readers of the archive, not as the archive. First of three companion works; the companions take up the questions deliberately left open.
Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. They are underspecified, they carry attachments, and they lack ground truth. We report K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs. Three blinded language-model judges scored every run against an eight-dimension rubric. On a rubric whose 8-anchor is defined as work a domain scientist would accept with minor edits, no model clears the line under all three judges. gpt-5.6-sol has the highest pooled mean, 8.04, but its 95% interval [7.80, 8.23] spans the threshold, and two of the three judges rank claude-opus-5 first instead. We therefore report the ordering of systems as the reproducible quantity, the absolute level as an attribute of the instrument, and the top of the table as unresolved. Across all 39,934 scored judgments -- the eight dimension scores plus a holistic overall for each assessment, excluding not-applicable cells -- 47.6% fall below the 8-point threshold. Difficulty is not uniform across the rubric: scientific accuracy averages 6.22 against 7.33 for communication, on identical denominators and in the same direction within every one of the nine models. The single leading failure tag is overclaiming, on 31.4% of assessments. We argue that the informative quantity for scientific agents is not a leaderboard position but the joint distribution of what was delivered, what was claimed, and what artifacts were produced.
Aubrey M. Brueckner, D. Patel, Yuhuan He et al.· 1 citation
Reach audiences
Advertise in front of researchers, engineers, and readers.
Comprehensively evaluating AI agents across interactive environments is difficult due to fragmented tasks, scaffolds, verifiers, and scoring rules. Unfortunately, existing efforts to unify these evaluations are limited in scale and domain, making costly reruns necessary and leaving available data incomparable. We introduce MESSIER, a unified corpus of 957,611 records spanning 30 benchmarks, 745 agents, 11,891 tasks, and 74,263 verifiers. MESSIER combines public evaluation results with new runs on six underrepresented professional and scientific benchmarks, standardizing their heterogeneous components into a common schema. Using this corpus, we show that frontier progress is uneven across benchmark groups, with function-calling evaluations largely saturated, programming improving fastest, and enterprise workflows remaining most challenging. Counterfactual rescoring further shows that strict all-pass scoring in multi-verifier tasks can alter agent rankings. Finally, we derive capability scores from our corpus that correlate with Epoch's Evaluation Capability Index rankings at Spearman \r{ho} = 0.84. The scores can also be estimated for subsets defined by domain, occupation, action space, or verifier type. In essence, MESSIER is a reusable resource for studying agent performance at scale, and a basis for designing better evaluations.
Stefan Krsteski, Charlotte Meyer, Guillaume Allegre et al.· 0 citations
Construction firms operate in knowledge-intensive and complex project environments where critical project knowledge is frequently fragmented across teams, documents, and digital systems. This fragmentation limits the systematic capture, structuring, and reuse of knowledge, which are fundamental processes for reducing knowledge loss, improving decision making, and enhancing project performance across organizational boundaries. Although effective knowledge mapping (KMp) is recognized as a valuable mechanism for organizing and disseminating such information, empirical evidence remains limited regarding how organizational, human, and technological factors interact to influence its effectiveness in construction contexts. This study addresses this gap by examining the interrelationships among these factors and their collective impact on KMp within construction firms. A quantitative, survey-based methodology was used to gather data from professionals within the construction industry. Structural equation modeling (SEM) using SPSS 23 and AMOS 24 was used to analyze the data and validate the measurement constructs and model fit. The study found that organizational frameworks, technological infrastructures, and human competencies significantly influence the effectiveness of KMp. Technological advancements were identified as a mediating factor, emphasizing the need for integration between digital tools and organizational culture to enhance knowledge-sharing processes. This study contributes to the knowledge management field by providing a systematic, data-driven perspective on the enablers of effective KMp. It extends the discourse on digital transformation and knowledge management, highlighting the importance of sociotechnical alignment between workforce capabilities and technological infrastructures. Future research should explore the role of artificial intelligence and machine learning in automating KMp processes. Practically, the findings provide construction firms with a structured strategy that integrates organizational alignment, workforce development, and technological investment to enhance KMp effectiveness and improve project decision making.
Safi Ullah, Xiaopeng Deng, Diana R. Anbar· Journal of construction engi...· 0 citations
A domain-specific legal artificial intelligence system for construction contract disputes via hybrid knowledge integration based on the retrieval-augmented generation (RAG) paradigm, integrating five core legal texts and 500 adjudication cases within a dual-engine architecture is proposed.
Ying Lu, Xin-Yun Shen, Yujing Wang et al.· Journal of construction engi...· 0 citations
Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value. Existing evaluations typically report performance metrics, but rarely examine agentic viability: whether dynamic LLM-mediated decisions convert their induced costs into measurable incremental profit. To apply this criterion, we introduce TradeLens, a trace-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations. It reconstructs trading trajectories, attributes profit and cost to interpretable evidence, and diagnoses whether and why an agent pays for its own intelligence. We conduct extensive analysis across backbone models, capital scales, trading frequencies, and system architectures, together with deployment discussion. Our results show that viability hinges on intelligence-to-profit conversion: models exhibit different failure patterns, such as poor asset selection in DeepSeek-V3.2 and negative timing in GLM-4.7, while capital scale, trading frequency, and architecture matter only by amplifying or degrading decision-attributed timing value. These findings reframe the evaluation of LLM-based trading agents from capability-centric performance ranking to trace-grounded diagnosis of intelligence-to-profit conversion. Our code is available at https://anonymous.4open.science/r/TradeLens.
Qiqi Duan, Changlun Li, Chen Wang et al.· 0 citations
Covalent inhibitors have garnered renewed attention in recent years, with their rational design becoming increasingly critical in drug discovery. Among the technologies facilitating the discovery of covalent inhibitors, covalent docking has emerged as a pivotal tool in various stages of drug development including virtual screening, lead optimization, and mechanistic studies. Since its inception as an extension of conventional docking methods in the early 2000s, covalent docking tools have undergone substantial advancements. This review provides a comprehensive overview of covalent docking algorithms, systematically categorizing their approaches according to covalent bond formation, which primarily include tethered docking, biased docking, and dynamic covalent docking approaches. A comparative analysis of current covalent docking tools is provided, alongside a critical discussion of remaining challenges. Special emphasis is placed on the growing impact of artificial intelligence (AI) in shaping novel methodologies and expanding the capabilities of covalent docking. Finally, we discuss prospects for advancing covalent docking methodologies and their applications in drug discovery.
Shi Li, Hongyan Du, Hui Zhang et al.· WIREs Computational Molecula...· 2 citations
Atomic charge is a fundamental quantum chemical property essential for advancing drug design and discovery. Although quantum mechanics (QM) methods offer the highest level of accuracy, their computational demands scale quadratically with the number of atoms, limiting their practicality for large-scale applications. In light of this, empirical and semiempirical methods have been introduced to improve computational efficiency, albeit often at the expense of accuracy. The advent of artificial intelligence has witnessed a growing application of machine learning (ML) techniques to accelerate atomic charge predictions. However, existing ML models often suffer from low accuracy and limited generalization capabilities. To address these challenges, we introduce an advanced equivariant graph attention neural network specifically engineered to model long-range atomic electrostatic interactions with high precision. This model introduces a sophisticated global graph attention mechanism, enabling it to capture charge contributions across multiple scales. By utilizing a combination of structural symmetry-preserving transformations and multiscale attention, our approach not only preserves the inherent symmetries of molecular structures but also substantially improves the model's accuracy, generalization, and robustness in complex scenarios. Our empirical analyses demonstrate that, compared to leading baseline models, the proposed model improves charge prediction accuracy by over 40% on average across various charge-calculation schemes. Remarkably, the model achieves superior performance on the external RESP (restrained electrostatic potential) test data sets, with a 54.6% improvement over the baseline. Additionally, we evaluated our charge model under the setting of virtual screening, where it outperforms both the OPLS3 charges and baseline deep learning models across all evaluation metrics, highlighting its extensive potential for scientific discovery.
Qiaolin Gou, Qun Su, Jike Wang et al.· Journal of Chemical Informat...· 1 citation
Colorectal cancer (CRC) exhibits substantial molecular heterogeneity, necessitating the inference of subtype-specific driver genes and their interactions for drug-target discovery and precision oncology. Prior studies often fail to capture subtle, latent nonlinear regulatory mechanisms (dark causal relationships) driving tumor progression in specific subtypes. Here, we develop an explainable intelligence computational framework, Symbolic Trajectory-Embedded Dark Causal Interaction Inference (STE-DC2I), which combines symbolic trajectory embedding with historical prediction mechanisms to model nonmonotonic oscillatory dependencies between genes. Integrating single-cell transcriptomic and multiomics profiles from malignant epithelial subpopulations, STE-DC2I classifies CRC subtypes, reconstructs developmental trajectories, and uncovers interpretable subtype-specific driver genes with functional relevance. Unlike correlation-based and explicit causal approaches, STE-DC2I captures weak yet biologically critical regulatory signals, outperforming state-of-the-art methods in predicting subtype-specific CRC driver genes. Functional assays in CRC cell lines (in vitro) validated nine predicted driver genes, highlighting their therapeutic potential.This work systematically explores dark causal interactions between genes in CRC subtypes. STE-DC2I offers interpretable insights and a generalizable strategy for CRC drug-target discovery.
Meng Huang, Huijin Hu, Ming Li et al.· Journal of Chemical Informat...· 0 citations
ConspectusThe field of covalent drug discovery has witnessed a remarkable resurgence in recent years, a trend underscored by the approval of more than 125 covalent drugs by the US FDA as of 2025, which demonstrates their immense therapeutic potential. Driven by ever-increasing computational power and vast amounts of data, deep learning (DL) is profoundly transforming numerous fields, from natural language processing to drug discovery. In the development of covalent drugs, in particular, advanced computational methods centered on data-driven approaches and artificial intelligence (AI) exhibit immense potential. The realization of this potential depends on the construction of a synergistic ecosystem. Here, we define this "ecosystem" as an integrated set of components─including (i) curated covalent-relevant databases, (ii) AI/physics-based predictive and scoring models, (iii) interoperable computational workflows spanning site identification, docking/virtual screening, and lead optimization, and (iv) closed-loop feedback that systematically incorporates experimental outcomes to update data resources and refine/validate models. This begins with the systematic collection of past experimental results to build high-quality databases. These databases, in turn, provide the foundation for developing AI-driven computational tools capable of precisely interfacing with and accelerating downstream tasks, such as molecular docking (for generating physically plausible conformations and conducting large-scale virtual screening) and lead optimization. The application of these AI tools not only guides experimental design, but the resulting key data also feed back into and enrich the databases. Furthermore, in the cutting-edge field of covalent drugs, the precise identification of "druggable" covalent sites on target proteins has emerged as another critically important downstream task.In this Account, we describe a computational and AI-driven ecosystem for structure-based covalent drug discovery and highlight our contributions to this field. By explicitly linking databases, models, workflows, and experimental feedback into a single framework, this Account moves beyond a simple inventory of individual tools to instead offer a systematic and panoramic perspective on an integrated ecosystem for covalent drug discovery, driven by data and computational engines including AI. We focus on how this ecosystem systematically addresses the challenges from covalent binding site identification to lead discovery, thereby fundamentally accelerating the development of next-generation covalent therapies. We first articulate the philosophy behind the construction and updating of covalent databases, emphasizing the necessity of high-quality data. Subsequently, we delve into a suite of cutting-edge, AI-driven computational methods, exploring the potential of deep learning in tasks such as molecular docking, covalent binding site prediction, and lead optimization. To bridge the gap between computational theory and experimental validation, we will use the discovery of potent covalent CRM1 inhibitors as a specific case study, detailing how our customized, structure-based virtual screening pipeline was utilized to achieve a seamless workflow from computational prediction to biological validation. This section is intended to offer actionable guidance for experimental researchers seeking to leverage these powerful computational tools. Finally, we highlight the limitations and potential pitfalls of this AI engine─concerns that are equally relevant when developing AI-driven covalent docking algorithms. Building on our group's recent benchmarking of AI docking methods, we objectively evaluate current performance and discuss how transformative advances such as AlphaFold3 may reshape the field.
Shi Li, Hongyan Du, Xujun Zhang et al.· Accounts of Chemical Researc...· 4 citations