Dynamic programming (DP) yields exact quadratic-time (O(NM)) pairwise sequence alignments. Static banding heuristics (O(NW)) fail catastrophically on low-identity (< 30%), asymmetric insertions/deletions (indels), or extreme length ratios, dropping core-block Sum-of-Pairs (SP) score recovery to 20%– 50%. Conversely, recent protein language model (PLM) aligners evaluate all N × M cells without search grid constraints. To bridge this gap, we introduce Adaptive-Banding Needleman–Wunsch (AB-NW), leveraging PLM contextual representations to construct a confidence-adaptive DP corridor prior to fine-resolution DP while keeping downstream scoring unmodified. AB-NW downsamples residue embeddings, computes a coarse alignment, and sets per-row corridor bounds via normalized confidence metrics. Evaluated via JIT-compiled buffers, this reduces time complexity to and space to O(NWmax) . Benchmarked across three PLM backbones (ESM2-8M, ESM2-35M, ProtBERT) across nine structural challenge categories, AB-NW recovers > 98.9% of exact unconstrained alignment scores and core-block SP accuracy across static banding failure modes (Twilight Zone, Asymmetric Indels, Extreme Aspect Ratios) while eliminating 55.3%–78.8% of active DP cells. On large protein matrices (N, M ≥ 3, 700), AB-NW eliminates 87.6%– 91.7% of cells, achieving speedups of 9.79×–13.30× (pure DP) and 1.73×– 2.94× (end-to-end), reaching up to 18.12× on unbiased controls (p < 0.05 to p < 10−15), making AB-NW practical for large-scale, high-throughput sequence alignment pipelines.
Experiments show that AP-REASONER outperforms baseline subsamplers on structure-sensitive downstream tasks and enables controllable recovery of alternative protein conformations, highlighting the value of modeling MSA subsampling as a controllable optimization problem, where factor-graph reasoning offers an effective a...
Proteins perform diverse cellular functions, and even single amino-acid substitutions can alter stability, activity, or molecular interactions. Protein language models (PLMs) provide a scalable approach for modeling such sequence--function relationships from unlabeled sequences, but increasing the size of dense Transfo...
Ming-Rui Li, Si-Xian Shen, Min-Zhang Li et al.· 0 citations
SNIPER is presented, a two-stage structured pruning framework that solves a knapsack optimization over coarse-granularity components to yield conditionally optimal parameter allocations with respect to fixed importance estimates, followed by a fine-grained pruning stage to meet strict budget constraints.
Palaash Goel, Ayan Sengupta, A. Nambi et al.· 0 citations
The results establish a passive, data-free provenance signal for compatible open-weight language-model checkpoints, and the projection-pairing signal appears across six language-model families and beyond.
It is discovered that outlier dominance grows with model scale: full replacement at K=128 degrades GPT-2 small but catastrophically degrades Mistral-7B by +1,134,279%, while selective replacement at p=99% rescues both models to under +15%.
Joseph Bingham· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.