Skip to content
#protein folding Dataset Open access

LIVIA Atlas: AlphaFold-Multimer protein interaction screens resolved to residues

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Data behind LIVIA Atlas: AlphaFold-Multimer protein interaction screens scored with iLIS (AFM-LIS), keeping the interface and contact residues of every prediction. livia-atlas-human-proteome.zip: human proteome-wide screen. Schmid et al. (2025), doi:10.1101/2025.11.10.687652. livia-atlas-human-kinase-tf.zip: human kinase–transcription factor screen. Kim et al. (2025), doi:10.1101/2025.10.10.681672. livia-atlas-flypredictome.zip: FlyPredictome, Drosophila melanogaster. Kim et al. (2026), doi:10.64898/2026.04.14.718529. Keyed by FlyBase gene; identity.tsv maps every construct name to its gene, and sets.json marks subsets such as the fly kinase–TF screen (Kim et al. 2025). livia-atlas-{human,zebrafish,yeast,worm}-kinase-kinase.zip: kinase–kinase screens in human, zebrafish, yeast and C. elegans. livia-atlas-viral-dimers-afdb.zip: protein pairs of 2,765 viruses (1,699,288 predictions: every pair within each viral proteome, homodimers included, never across viruses) from the AlphaFold Database viral protein complexes release (EMBL-EBI, Google DeepMind, NVIDIA and collaborators, 2026; CC BY 4.0), one AlphaFold-Multimer model per pair, rescored with lis.py. livia-atlas-afdb-heterodimers-part1.zip to part6.zip: heterodimer screens from the AlphaFold Database protein complexes released with NVIDIA (EMBL-EBI, NVIDIA and collaborators; CC BY 4.0; ftp.ebi.ac.uk/pub/databases/alphafold/collaborations/nvda/heterodimers/), one screen per species (7,573,187 protein pairs, 312,636 proteins), rescored with lis.py. The screens are grouped into six zips because a Zenodo record holds at most 100 files.livia-atlas-afdb-heterodimers-structure-index.zip: where each heterodimer model and its PAE sit in the release archive at EBI, so a single model can be read by byte range. livia-atlas-heterodimers-lis-part01.csv.zst and part02.csv.zst: every heterodimer pair as lis.py scored it (all columns), with the address of its model in the EBI archive. Zstandard-compressed; README_heterodimers_lis_table.md explains how to read the table and open one structure. Each zip is uncompressed, so the website reads single files by byte range. Inside: manifest.json, proteins.json, edges.tsv (pairs past the 10% FPR cutoff), and one bundle per protein in b/ (lis.py rows with residues, and a FASTA). Please cite the source of each screen. Version 0.1.1: FlyPredictome constructs carry the sequences that were folded and are placed on their genes by sequence; kinase–kinase screens added. Version 0.1.2: a gene whose isoforms were folded separately keeps one bundle per isoform, so a page can read the reference first: b/ .zip holds the reference and an isoforms.json listing the others, each in b/ ~ .zip. Version 0.1.3: the viral screen is added. In the existing screens, scores are cut, not rounded (4 decimals in bundles, 3 in edge lists); 3,192 residue lists that the old scorer had shortened are emptied and flagged; missing scores are null; a sequence pair folded more than once is counted once, by its highest-iLIS prediction, while each screen keeps its own; and the fly kinase–kinase screen is complete (34,271 pairs; one batch had been filed with the kinase–TF screen). Version 0.1.4: the AlphaFold Database heterodimer screens are added, with the full lis.py tables and a structure index. The files of 0.1.3 are unchanged (same checksums). Versions were renumbered from 1.x to 0.1.x on 27 September 2026 (1.0 is 0.1.0, 1.1 is 0.1.1, 1.2 is 0.1.2); their files and DOIs are unchanged.

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or seque...

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.