Skip to content
← All posts When AI art has no author: Study finds generated images often can’t be traced to training data

When AI art has no author: Study finds generated images often can’t be traced to training data

MIT News · Artificial Intelligence · news.mit.edu · By Rachel Gordon | MIT CSAIL · August 18, 2026

A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.

Read on MIT News · Artificial Intelligence → Opens the original article in a new tab.

More from the blog

Related papers

#artificial intelligence Preprint Jul 2026

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitutional corpus from Anthropic's Constitution and apply constitutional midtraining at 120B scale, where principled, values-based content is inserted into midtraining. A 2x2 design (curriculum ordering x deliberative reasoning) was used to produce four constitutionally midtrained conditions, plus a control, which were evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment. All models were evaluated across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperformed the control on alignment generalization and durability, notably on blackmail: SFT instilled a blackmail propensity in all models, but constitutional midtraining blunted it, with the advantage surviving benign fine-tuning (-17.5pp). This durability did not extend to settings that required active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also mattered more than its structure, and constitutional midtraining incurred no capability cost, on average, at any stage (MMLU, ARC-Easy, piqa, GSM8K). A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.

Desiree Cho, Cameron Tice, Bernie Hogan et al. · 0 citations
#artificial intelligence Review Aug 2026

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model responses to augment opinion survey data, our procedure achieves target coverage and substantially narrows confidence intervals.

Chengpiao Huang, Kaizheng Wang · 0 citations
#diffusion models Open access Aug 2026

Application of Artificial Intelligence Frameworks in Development of Medical Imaging Diagnosis Systems: A Comprehensive Review of Novel Methodologies, Clinical Validation, Performance Optimization, and Future Perspectives

Artificial intelligence (AI) has revolutionized medical imaging with automated disease detection, image segmentation, diagnosis, prognosis prediction, and clinical decision support across various imaging modalities, such as X-ray, computed tomography (CT), magnetic resonance imaging (MRI), ultrasound, positron emission tomography (PET), retinal imaging, and digital pathology. Recent advances in deep learning, transformer architectures, multimodal learning and foundation models have significantly improved the diagnostic accuracy and reduced the reliance on handcrafted feature engineering. However, challenges like data heterogeneity, model interpretability, external validation, privacy preservation, computational efficiency, and regulatory compliance still hinder the widespread clinical implementation.In this paper, this review presents a comprehensive study on the evolution of modern AI frameworks in medical imaging by integrating recent methodological advances with perspectives on clinical translation. The review covers the latest deep learning architectures, such as convolutional neural networks, Vision Transformers, hybrid CNN–Transformer models, multimodal learning frameworks, generative artificial intelligence, diffusion models, federated learning, privacy-preserving learning, and medical foundation models. In addition, the review covers the cutting-edge explainable AI techniques, including Grad-CAM, SHAP, LIME, and attention visualization, for boosting transparency and clinician confidence. The review also discusses the state-of-the-art performance optimization strategies, including transfer learning, active learning, domain adaptation, neural architecture search, hyperparameter optimization, model compression, and computational resource optimization. Equally important, the latest developments in clinical validation, external evaluation, robustness assessment, fairness, uncertainty estimation, regulatory considerations, and deployment frameworks are critically analyzed to underscore their role in facilitating safe clinical implementation.The review analysis concludes with the identification of key research challenges and directions for the future including multimodal foundation models, vision-language systems, retrieval-augmented generation,

S Sur, Mandal Rakesh Kumar, Chanda Debanil · 0 citations
#diffusion models Open access Aug 2026

Application of Artificial Intelligence Frameworks in Development of Medical Imaging Diagnosis Systems: A Comprehensive Review of Novel Methodologies, Clinical Validation, Performance Optimization, and Future Perspectives

Artificial intelligence (AI) has revolutionized medical imaging with automated disease detection, image segmentation, diagnosis, prognosis prediction, and clinical decision support across various imaging modalities, such as X-ray, computed tomography (CT), magnetic resonance imaging (MRI), ultrasound, positron emission tomography (PET), retinal imaging, and digital pathology. Recent advances in deep learning, transformer architectures, multimodal learning and foundation models have significantly improved the diagnostic accuracy and reduced the reliance on handcrafted feature engineering. However, challenges like data heterogeneity, model interpretability, external validation, privacy preservation, computational efficiency, and regulatory compliance still hinder the widespread clinical implementation.In this paper, this review presents a comprehensive study on the evolution of modern AI frameworks in medical imaging by integrating recent methodological advances with perspectives on clinical translation. The review covers the latest deep learning architectures, such as convolutional neural networks, Vision Transformers, hybrid CNN–Transformer models, multimodal learning frameworks, generative artificial intelligence, diffusion models, federated learning, privacy-preserving learning, and medical foundation models. In addition, the review covers the cutting-edge explainable AI techniques, including Grad-CAM, SHAP, LIME, and attention visualization, for boosting transparency and clinician confidence. The review also discusses the state-of-the-art performance optimization strategies, including transfer learning, active learning, domain adaptation, neural architecture search, hyperparameter optimization, model compression, and computational resource optimization. Equally important, the latest developments in clinical validation, external evaluation, robustness assessment, fairness, uncertainty estimation, regulatory considerations, and deployment frameworks are critically analyzed to underscore their role in facilitating safe clinical implementation.The review analysis concludes with the identification of key research challenges and directions for the future including multimodal foundation models, vision-language systems, retrieval-augmented generation,

S Sur, Mandal Rakesh Kumar, Chanda Debanil · 0 citations