Skip to content
#protein folding Open access

Doulita Token Compressor: Geometric Context Compression for LLMs - Research Preview. Duqueana Core

Aug 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Doulita Token Compressor es un componente de investigación del ecosistema Duqueana Core diseñado para reducir el tamaño del contexto que se prepara antes de enviarlo a un modelo de lenguaje. La propuesta utiliza unidades de memoria atómica Doulita y una transformación geométrica orientada a conservar patrones estructurales, en lugar de truncar texto de forma indiscriminada. La demostración ejecutada que acompaña este informe generó 1.000 registros, con una entrada de 71.889 caracteres y una estimación de aproximadamente 17.972 tokens originales. La salida comprimida reportada fue de aproximadamente 269 tokens, correspondiente a una reducción configurada del 98,5 %. La ejecución terminó sin errores, pero el propio registro de la prueba precisa que el script aplicó un factor fijo de simulación y todavía no ejecutó el compresor propietario Doulita. Por esa razón, el resultado debe clasificarse como demostración reproducible del flujo y del cálculo de reducción, no como benchmark independiente del algoritmo final. Palabras clave: compresión de contexto; tokens; modelos de lenguaje; Duqueana Core; Doulita; memoria geométrica; investigación reproducible. { "@context": "https://schema.org", "@type": "SoftwareApplication", "name": "Duqueana Core", "description": "Framework avanzado de simulación impulsado por el motor MREI, capaz de ejecutar modelos complejos con alta eficiencia en hardware clásico.", "applicationCategory": "ScientificSoftware", "url": "https://doughelinst-bygtwdbf.manus.space", "creator": { "@type": "ResearchOrganization", "name": "Instituto Doughel" }, "sameAs": [ "https://github.com/dougheliano-beep/DUQUEANA--CORE-", "https://zenodo.org/communities/post-classical-computing/" ] } { "@context": "https://schema.org", "@type": "SoftwareApplication", "name": "Duqueana Core", "alternateName": "Framework de Computación Post-Clásica", "headline": "Framework de simulación y computación post-clásica basado en iteración de estados", "description": "Duqueana Core es un framework de simulación computacional post-clásica basado en el motor MREI (Resolución Iterativa de Estados). Implementa principios de la IA Duqueana, un paradigma post-IA que prioriza estructura, determinismo y eficiencia sobre el aprendizaje estadístico. Incluye módulos de simulación científica, reducción extrema de tokens y un asistente conversacional estructural. Probado en múltiples entornos de IA para validar estabilidad, compatibilidad y reproducibilidad.", "applicationCategory": "ScientificSoftware", "softwareVersion": "2.1", "operatingSystem": "Cross-platform", "featureList": [ "Simulación determinista mediante iteración de estados (MREI)", "Paradigma post-IA basado en estructura y no en datos", "Reducción extrema de tokens (90–98%)", "Interpretación conversacional estructural (ACD)", "Ejecución eficiente en hardware clásico", "Modelado bioquímico y físico reproducible", "Arquitectura verificable y auditable", "Probado en múltiples entornos de IA (Qwen, DeepSeek, Minimax, Copilot, Manus)" ], "testingInformation": { "@type": "CreativeWork", "name": "Validación en entornos de IA", "description": "Duqueana Core ha sido probado en diferentes modelos y plataformas de IA para evaluar estabilidad, compatibilidad y comportamiento estructural. Las pruebas incluyeron Qwen, DeepSeek, Minimax, Copilot y Manus, confirmando la reproducibilidad del motor MREI y la eficiencia del módulo de reducción de tokens." }, "creator": { "@type": "ResearchOrganization", "name": "Instituto Doughel de Investigación Digital", "url": "https://doughelinst-bygtwdbf.manus.space" }, "author": { "@type": "Person", "name": "Douglas Helvesio Urbina Duque", "affiliation": "Instituto Doughel", "identifier": "ORCID: 0009-0005-1230-7549" }, "keywords": [ "Duqueana Core", "IA Duqueana", "post-IA", "computación post-clásica", "iteración de estados", "MREI Engine", "simulación científica", "reducción de tokens", "ACD", "paradigma computacional", "validación en IA", "Qwen", "DeepSeek", "Minimax", "Copilot", "Manus" ], "url": "https://doughelinst-bygtwdbf.manus.space", "codeRepository": "https://github.com/dougheliano-beep/DUQUEANA--CORE-", "sameAs": [ "https://zenodo.org/communities/post-classical-computing/", "https://zenodo.org/records/22168535" ], "license": "https://opensource.org/licenses/MIT" } { "@context": "https://schema.org", "@type": "SoftwareApplication", "name": "Duqueana Core", "alternateName": "Post-Classical Computing Framework", "headline": "Framework de simulación post-clásica con 98% de eficiencia energética", "description": "Duqueana Core es una plataforma de computación post-clásica que incluye: MREI Engine (simulación determinista), Doulita Token Compressor (90-98% reducción de tokens para LLMs), ACD (asistente conversacional estructural) y módulos de simulación bioquímica (p53, FeMo-Co). Responde a crisis de IA: opacidad, consumo energético desbordado y centralización. Ejecutable en hardware estándar.", "applicationCategory": "ScientificSoftware", "softwareVersion": "2.1", "operatingSystem": "Cross-platform", "featureList": [ "MREI Engine: simulación determinista mediante iteración geométrica", "Doulita Token Compressor: 90-98% reducción de contexto LLM", "ACD: asistente conversacional basado en estructura, no estadística", "Simulación bioquímica: p53, FeMo-Co, DNMT3A/B", "Simulación física: N-Body, circuitos cuánticos 53-qubit", "Eficiencia energética: 65-81% menos RAM que métodos clásicos", "Hardware estándar: sin dependencia de clusters de élite", "Verificabilidad: cada módulo con hashes SHA-256 y manifiestos", "Validación multi-IA: Qwen, DeepSeek, Minimax, Copilot, Manus" ], "creator": { "@type": "ResearchOrganization", "name": "Instituto Doughel de Investigación Digital", "url": "https://doughelinst-bygtwdbf.manus.space" }, "author": { "@type": "Person", "name": "Douglas Helvesio Urbina Duque", "affiliation": "Universidad Nacional Experimental de Guayana (UNEG)", "identifier": "ORCID: 0009-0005-1230-7549" }, "keywords": [ "Duqueana Core", "post-classical computing", "MREI Engine", "Doulita Token Compressor", "token compression 98%", "energy-efficient AI", "sustainable computing", "green AI", "AI carbon footprint reduction", "verifiable computing", "AI governance", "deterministic simulation", "structural intelligence", "p53 simulation", "FeMo-Co modeling", "N-Body simulation", "53-qubit circuit sampling", "standard hardware", "democratized AI", "ACD conversational assistant", "geometric iteration", "computational efficiency", "biochemical modeling", "quantum circuit simulation", "protein folding", "computational chemistry" ], "hasPart": [ { "@type": "SoftwareApplication", "name": "MREI Engine", "description": "Motor de Resolución Exacta Iterada para simulación determinista", "applicationCategory": "Simulation Engine" }, { "@type": "SoftwareApplication", "name": "Doulita Token Compressor", "description": "Compresor geométrico de contexto LLM: 90-98% reducción de tokens. Research Preview con DOI 10.5281/zenodo.22168535", "applicationCategory": "AI Efficiency Tool", "keywords": ["token reduction", "energy efficiency", "LLM optimization", "98.5% compression"] }, { "@type": "SoftwareApplication", "name": "ACD (Asistente Conversacional Duqueano)", "description": "Intérprete conversacional post-clásico basado en simulación estructural", "applicationCategory": "Conversational AI" }, { "@type": "CreativeWork", "name": "p53 Restoration Simulator", "description": "Simulación de reparación de mutaciones TP53 hotspot" }, { "@type": "CreativeWork", "name": "FeMo-Co Modeling", "description": "Simulación de cofactores metálicos y fijación de nitrógeno" }, { "@type": "CreativeWork", "name": "N-Body Gravitational Simulation", "description": "Dinámica gravitacional iterativa post-clásica" }, { "@type": "CreativeWork", "name": "53-Qubit Circuit Sampling", "description": "Replicación clásica del experimento Google Sycamore 2019" } ], "url": "https://doughelinst-bygtwdbf.manus.space", "codeRepository": "https://github.com/dougheliano-beep/DUQUEANA--CORE-", "sameAs": [ "https://zenodo.org/communities/post-classical-computing/", "https://zenodo.org/records/22168535", "https://zenodo.org/records/22119449" ], "license": "https://creativecommons.org/licenses/by-nc-nd/4.0/", "contributor": [ { "@type": "Organization", "name": "Universidad Nacional Experimental de Guayana (UNEG)", "description": "Validación académica en física y química" }, { "@type": "Organization", "name": "Universidad de los Andes Venezuela", "description": "Colaboración académica" }, { "@type": "Organization", "name": "Universidad del Zulia", "description": "Colaboración académica" } ], "audience": { "@type": "Audience", "audienceType": [ "Researchers", "Bioinformaticians", "AI Ethics Specialists", "Sustainability Engineers", "Post-Classical Computing

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

In the context of cloud computing, risks associated with underlying technologies, risks involving service models and outsourcing, and enterprise readiness have been recognized as potential barriers for the adoption. To accelerate cloud adoption, the concrete barriers negatively influencing the adoption decision need to be identified. Our study aims at understanding the impact of technical and security-related barriers on the organizational decision to adopt the cloud. We analyzed data collected through a web survey of 352 individuals working for enterprises consisting of decision makers as well as employees from other levels within an organization. The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability. The result from our logistic regression analysis confirms the criticality of the security concern, which results in an up to 26-fold increase in the non-adoption likelihood. Our study underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

To compete in this age of disruption, large companies cannot rely on cost efficiency, lead time reduction and quality improvement. They are now looking for ways to innovate like startups. Meanwhile, the awareness and use of the Lean startup approach have grown rapidly amongst the software startup community in recent years. This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors. A multiple case study approach is followed in the investigation. Two software product innovation projects from two large companies are examined, using a conceptual framework that is based on the method-in-action framework and extended with the previously developed Lean-Internal Corporate Venture model. Seven face-to-face in-depth interviews of the employees with different roles are conducted. Within-case analysis and cross-case comparison are applied to draw the findings from the cases. A generic process flow summarises the common key processes of Lean internal startups. The findings suggest that an internal startup that is initiated management or employees faces different challenges. A list of enablers of applying Lean startup in large companies are identified, including top management support and cross-functional team. Both cases face different inhibitors due to the different process of inception, objective of the team and type of the product. Our contributions are threefold. First, this study is one of the first attempt to investigate the use of Lean startup approach in large companies empirically. Second, the study shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context. The third is a general process of Lean internal startup and the evidence of the enablers and inhibitors of implementing it, which are both theory-informed and empirically grounded.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

Affects--emotions and moods--have an impact on cognitive processing activities and the working performance of individuals. It has been established that software development tasks are undertaken through cognitive processing activities. Therefore, we have proposed to employ psychology theory and measurements in software engineering (SE) research. We have called it "psychoempirical software engineering". However, we found out that existing SE research has often fallen into misconceptions about the affect of developers, lacking in background theory and how to successfully employ psychological measurements in studies. The contribution of this paper is threefold. (1) It highlights the challenges to conduct proper affect-related studies with psychology; (2) it provides a comprehensive literature review in affect theory; and (3) it proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

It is essential for startups to quickly experiment business ideas by building tangible prototypes and collecting user feedback on them. As prototyping is an inevitable part of learning for early stage software startups, how fast startups can learn depends on how fast they can prototype. Despite of the importance, there is a lack of research about prototyping in software startups. In this study, we aimed at understanding what are factors influencing different types of prototyping activities. We conducted a multiple case study on twenty European software startups. The results are two folds; firstly we propose a prototype-centric learning model in early stage software startups. Secondly, we identify factors occur as barriers but also facilitators for prototyping in early stage software startups. The factors are grouped into (1) artifacts, (2) team competence, (3) collaboration, (4) customer and (5) process dimensions. To speed up a startup’s progress at the early stage, it is important to incorporate the learning objective into a well-defined collaborative approach of prototyping.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Sep 2026

Silk fibroin-loaded Fe-curcumin nanoparticles on antimicrobial peptide-functionalized TiO2 nanotube surfaces: Microenvironment-modulated synergy for antibacterial and osteogenic enhancement.

This study constructed a pH-responsive P-TN/SF@Fe-Cur composite coating that demonstrated significant anti-infective, anti-inflammatory, antioxidant, pro-angiogenic, and pro-osteogenic effects in rat subcutaneous infection and femoral defect models.

Xiaotong Shen, Shuxia Huang, Ting Zhang et al. · 2 citations
#computer vision May 2016

Bringing the Cloud to Rural and Remote Areas - Cloudlet by Cloudlet

Instead of relying on huge and expensive data centers for rolling out cloudbased services to rural and remote areas, we propose a hardware platform based on small single-board computers. The role of these micro-data centers is twofold. On the one hand, they act as intermediaries between cloud services and clients, improving availability in the case of network or power outages. On the other hand, they run community-based services on local infrastructure. We illustrate how to build such a system without incurring high costs, high power consumption, or single points of failure. Additionally, we opt for a system that is extendable and scalable as well as easy to deploy, relying on an open design.

P. Abrahamsson, S. Helmer, Tosin Daniel Oyetoyan et al. · 2 citations

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.