OreoLook (formerly lixSearch), an open-source answer engine using automated browser agents and provider-routed LLM inference, is developed, which presents a three-layer caching architecture that deduplicates embedding computations across sessions.
Abstract
AI-powered search products such as ChatGPT search, Google's AI Overviews, and Perplexity provide LLM-synthesized answers grounded in live web results. We developed OreoLook (formerly lixSearch), an open-source answer engine using automated browser agents and provider-routed LLM inference. Its local search, caching, session-management, and embedding stack runs on commodity CPU hardware; answer synthesis is performed by a remote inference provider. As usage grew, sessions lost context, equivalent queries triggered redundant work, and URLs were repeatedly embedded across sessions. We present a three-layer caching architecture: (1) a Session Context Window maintaining a rolling window of recent messages in Redis with automatic overflow to Huffman-compressed disk archives; (2) a Semantic Query Cache catches rephrasings via cosine similarity on embedding vectors, eliminating redundant LLM invocations; and (3) a URL Embedding Cache that deduplicates embedding computations across sessions. Deployed on a single 8-vCPU Intel Cascade Lake server (2 GHz, 32 GB RAM) running 30 Hypercorn worker processes across three containerized replicas, the evaluated system reported an 89.3% aggregate Redis keyspace hit rate with 0.1 ms read latency and just 1.38 MB of memory overhead. A background LRU eviction daemon migrates idle sessions from Redis to disk and re-hydrates them on demand, enabling conversations that can be resumed hours or days later under the configured retention policy.
The cross-encoder study shows that thresholds do not transfer between embedding models, and LFU is the strongest simple default in this protocol; deployment decisions should first establish answer validity and then test sub-point policy differences with exact search.
Y. Kulkarni, Shubham Harkare, A. Babu· 0 citations
Modern web applications demand sustained low latency under workloads that shift across users, devices, sessions, and network conditions. Classical cache replacement policies such as Least Recently Used (LRU) and Least Frequently Used (LFU) treat every cached object identically and ignore the cost-of-miss heterogeneity...
Akshatha Madapura Anantharamu· World Journal of Advanced En...· 0 citations
We describe the integration of io_uring into Oracle Database's storage layer and the architectural decisions required to deploy it in a production multi-process RDBMS. Our design uses per-process ring contexts that eliminate inter-process synchronization, a shared buffer registration mechanism now part of the mainline...
R. Chowdhury, A. Shah, Margaret Susairaj et al.· 0 citations
Log-structured merge-tree (LSM-tree) key-value stores rely on caching to mitigate multi-component lookups and long tails, yet block and KP caches are prone to compactioninduced expiry, and all three cache types suffer from scan pollution under LRU eviction. We present AutoThermKV, a two-level in-memory architecture tha...
Yunfan Chi, E. Sha, Longshan Xu et al.· IEEE International Conferenc...· 0 citations
Fault-tolerant, distributed shared logs are a useful substrate for distributed applications, but existing designs coordinate shard servers, sequencers, and replicas over the network, so append and replay latency is dominated by message exchange. We present Borges, to our knowledge the first fault-tolerant shared log wh...
Hao-Wei Chen, Yi-Ming Xiang, Zhi-Peng Jia et al.· Proceedings of the ACM SIGOP...· 0 citations
Prefix caching, in which a serving engine reuses the key and value tensors of a shared prompt prefix across requests, is enabled by default in the major open-source stacks and treated as a transparent optimization. We measure what it costs in reproducibility, and find that the cost rises sharply with weight quantizatio...
Aditi Patodiya· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.