Skip to content
Preprint

Scalable PII Discovery in Mobile App Databases via Hypothesis-Driven Search

Aug 2026 · 1 citation · 24 references
Computer Science

TL;DR

A hypothesis-driven framework that treats PII localization as bounded, adaptive search under uncertainty, and separates lightweight PII exploration from targeted extraction, normalization, and deduplication over validated regions, thereby limiting exhaustive inspection to regions supported by sampled evidence.

Abstract

Discovering personally identifiable information (PII) in mobile forensic databases is difficult because the relevant table-column regions are unknown, distributed across heterogeneous SQLite schemas, and may contain values embedded in free-text or semi-structured fields. We present a hypothesis-driven framework that treats PII localization as bounded, adaptive search under uncertainty. An agent ranks candidate table-column regions, probes sampled values, and maintains a memory of prior evidence, confidence scores, and decisions to refine subsequent hypotheses. The framework separates lightweight PII exploration from targeted extraction, normalization, and deduplication over validated regions, thereby limiting exhaustive inspection to regions supported by sampled evidence. We evaluate the framework on 25 SQLite databases from 10 Android and iOS applications in the Cellebrite CTF corpus, targeting email addresses, phone numbers, domain names, person names, and postal addresses. Against a corpus-level distinct ground-truth set of 3,751 entities, Gemini 2.5 Pro achieves 94.5% F1 while reducing the effective extraction search space by 79.9% on average. Results across 12 model backends show strong performance among several frontier models, but substantial sensitivity to model capability.

View source

Similar papers

Review Aug 2026

Security and Privacy Taxonomy Generation from Mobile App Reviews

This work filters app reviews for privacy- and security-related content, yielding a comprehensive corpus of over 600K reviews, and introduces TaxoScale, a pipeline that handles taxonomy construction at this scale by extending an expert-defined taxonomy via Recursive Hierarchical Clustering and LLM-based node naming.

Moghis Fereidouni, Vinaik Chhetri, Umar Farooq et al. · 0 citations
Preprint Aug 2026

The"Curse of Knowledge"in LLM Query Simulation: Concept Provenance for Tracing Answer-Side Intrusion

LLM-generated search queries are widely used to augment IR evaluation, yet they may contain concepts that presuppose answer-side document knowledge, violating the information-access boundary of pre-search users. Existing validation metrics, including overlap, diversity, and effectiveness, cannot distinguish rare human-...

Chenglong Ma, Xinye Wanyan, Danula Hettiachchi et al. · 0 citations
Preprint Aug 2026

Lost in Permissions: Exploring the Microsoft 365 App Ecosystem

This work presents the first privacy- and security-oriented measurement of M365 third-party applications, and finds that only 1,069 of them expose both descriptions and permission sets, with significant inconsistencies in transparency across official distribution channels.

Vincenzo Longo, Alberto Verna, Nikhil Jha et al. · 0 citations
Book Jul 2026

Building a User Foundation Model for the Open Web

User foundation models have demonstrated strong results in e-commerce and social recommendation, but most industrial deployments assume environments where user identity is stable and persistent. Open-web real-time bidding (RTB) operates on a structurally different data distribution: user identity is fragmented and non-...

Solal Vernier, Ivan Can Arisoy, Merwan Barlier et al. · 0 citations
Preprint Aug 2026

RecGPT-Mobile-V2 Technical Report

RecGPT-Mobile-V2 is introduced, an end-to-end framework that treats intent quality and execution efficiency as coupled objectives within a staged design and helps retain decision-relevant evidence and allocate additional computation only when it is likely to improve the predicted Query.

Lingqin Zhang, Bin Zhang, Wei-Peng Huang et al. · 0 citations
Preprint Aug 2026

Guardian Crawler: Retrieval-First Knowledge Discovery with Bounded LLM Augmentation for Noisy Web Intelligence

Retrieving relevant evidence from noisy web data is challenging, particularly in sensitive domains containing incomplete reports, heterogeneous language, and irrelevant content. We present Guardian Crawler, a reproducible retrieval-first testbed for controlled experiments on knowledge discovery and evidence-grounded su...

J. Castillo, S. Nukavarapu, Ravi Mukkamala · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.