Preprint
Jul 2026
Hollywood: Towards a Large Movie Dataset for Database Benchmarking
Hollywood is introduced, a synthetic IMDb-compatible benchmark generator that combines LLM-generated semantic dictionaries with deterministic temporal-graph-based relational data generation and demonstrates that Hollywood induces cardinality estimation errors comparable to or exceeding those observed on the original IMDb dataset.
Ivan Iachnyk, Mihail Stoian, Andreas Kipf
· 0 citations