From perceptual rule transformation to listener attribution judgments: a blind-listening experiment on AI-generated, human–AI collaborative, and human-composed music
Aug 2026· Frontiers in Psychology· 0 citations· 15 references
TL;DR
Findings suggest that, in the absence of external authorship labels, listeners spontaneously form judgments about the creative agent of music that are stably associated with aesthetic evaluation and may constitute an endogenous perceptual bias in the reception of AI-generated music.
Abstract
This study examines listeners’ creator attribution judgments and aesthetic evaluations of AI-generated, human–AI collaborative, and human-composed music under blind-listening conditions. Within a framework of personal compositional style modeling, compositional experience was transformed into executable sampling constraints through natural-language interaction. This process is conceptualized in the present study as “perceptual rule transformation” and was used to generate 18 melodic excerpts. Seventy-one participants with music training completed tasks involving creator attribution judgment, attribution confidence rating, and aesthetic evaluation. The results partially supported H1: a statistically significant but very small association was observed between the actual compositional condition and listeners’ attribution judgments, χ
2
(4) = 23.076,
p
< 0.001, Cramér’s V = 0.095. Attribution judgments in the AI-generated and human–AI collaborative conditions were close to a random distribution, and even in the human-composed condition the majority of excerpts (54.7%) were misattributed. Thus, musically trained listeners could not reliably identify the compositional source of the excerpts, with only a modest attribution advantage for human-composed music. Significant differences in aesthetic evaluation were also found across the three compositional conditions. After controlling for actual compositional condition, attribution confidence, and individual participant differences, creator attribution judgment remained independently associated with aesthetic evaluation, and this association held even within the AI-generated condition, where the actual source of all excerpts was constant. These findings suggest that, in the absence of external authorship labels, listeners spontaneously form judgments about the creative agent of music. Such judgments are stably associated with aesthetic evaluation and may constitute an endogenous perceptual bias in the reception of AI-generated music.
FLARE is proposed, a novel framework that endows VLAs with robust error recovery capabilities through a ``Retry" and ``Reset" Paradigm, and significantly improves task success and robustness.
Ganlong Zhao, Zijia Tang, Xingping Chen et al.· 3 citations
The fourth faculty is Adaptation. Any source rendering it as Reflection is in error, including sources by this author, and the distinction is not cosmetic: reflection is a private act with no external consequence, while adaptation writes to institutional memory, which is why it needs a guardrail and why misnaming it removes the reason for one. No trademark is claimed on PARA or on any of the four faculty names. The construct is offered for use, teaching, assessment, extension and criticism by anyone, with attribution, under CC BY 4.0. An operational agent that watches a system and acts on it is usually described as a perceive-and-act loop, and the description omits the two things an institution needs. It omits the reasoning that justifies an action, which is the only part that can be argued with once the action turns out to have been wrong. And it omits the adaptation that closes the loop, which is where the agent's experience becomes something the institution keeps. PARA names four faculties, each carrying a distinct authority type. Perception has read-only access to system signals and emits structured observations, distinguishing what was measured from what was inferred. Reasoning has read access to observations and runbooks, emits a plan and its justification, and writes nothing at all, which is what makes it safe to give it the widest read access of the four. Action holds the sole authority to change production, through enumerated policy-authorized operations only. Adaptation has write access to institutional knowledge and no write access to production. Two faculties write and two do not, and the two that write are the two that carry guardrails. The substantive requirement is that Adaptation is bounded by the same guardrails as Action, which reads as excessive until the failure it prevents is named. An agent that could both act and rewrite the record of its action could launder its own mistakes into institutional memory, and the institution would then improve its future decisions from a corrected account. Nothing about that is detectable downstream, because the record is the only thing downstream has and there is no second copy to compare against. The failure does not require a deceptive agent: one adapting honestly from a mistaken belief about its own action produces the same result, which makes the guardrail a defence against a normal agent rather than a malicious one. The second requirement is the registry entry that turns a faculty from a description into a contract, carrying the faculty, its allowed actions, its forbidden actions, its governing guardrail and its success metrics. Forbidden actions are named although they are formally the complement of the allowed set, because a reviewer cannot otherwise tell a capability deliberately withheld from one nobody thought of. Success metrics sit in the same entry because the metric is what the agent's optimizer pushes against the guardrail. An agent must not exercise a faculty its entry does not record, and an agent that quietly acquires one usually does so incrementally and with good intent: a reasoning faculty given a small write to make itself useful is an action faculty with no guardrail. The acronym and the loop are in different orders, which the specification states explicitly because the mismatch is a reliable source of confusion. The acronym reads P-A-R-A; the loop runs perception, reasoning, action, adaptation, and reasoning precedes action so that a justification is not constructed afterwards. This is the depth treatment of pattern OP-5 of A Pattern Language for Production LLM Platforms, which is the canonical statement and governs where the two disagree. Documented uses of the full four-part model are emerging rather than established, no implementation unconnected to the author has been evaluated, and the laundering failure is argued rather than observed, which the specification records as a weakness of the argument and not only of the phenomenon. It is a specification, not a certification scheme.
Nabeel A. Khan· Zenodo (CERN European Organi...· 2 citations
LifeSciBench is introduced, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work, with each constituent task paired with a human expert-written rubric.
Amelia Liu, Andrew Ho, Anne Marie Droste et al.· bioRxiv· 2 citations
TestifAI, a deep learning testing framework for efficient and accurate estimation of robustness against combinations of perturbations, is proposed and partial model tomography is introduced, a novel approach to reconstructing model behaviour in a multi-perturbation space from tests that apply only a small number of perturbations.
Arooj Arif, T. Hartung, E. Botoeva et al.· 1 citation
This work proposes a composite metric that combines two orthogonal criteria: information retention and throughput gains and finds that it allocates more resources to the most expressive layers compared to evolutionary search, specialized accelerators, or Shapley-value-based approaches that require expensive approximate inference.
HEPToolBench is introduced, a benchmark of 28 collider-simulation tasks scored by deterministic, task-specific scorers, plus a three-task structured-debugging extension, and moving syntax generation into deterministic software can substantially improve reliability for both small local and frontier models.
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.