FaaSHive: Addressing Limited Visibility in Function-as-a-Service via Worker-driven Scheduling
Abstract
Effective request placement in Function-as-a-Service (FaaS) platforms requires timely visibility into worker state, which changes rapidly as workers create, reuse, and evict short-lived function instances. Under the conventional cloud manager–worker architecture commonly used in FaaS platforms, however, this visibility is difficult to sustain because the central manager must make placement decisions using a cluster-level view of worker state. Even modest delays between state collection and scheduling can make this view stale before placement decisions take effect. The resulting mismatch between the manager’s view and the workers’ actual execution state is amplified by high consolidation in FaaS deployments, leading to inaccurate placement, additional cold starts, lower throughput, and higher latency. We present FaaSHive, a worker-driven FaaS scheduling framework that mitigates this visibility gap by partitioning workers into groups and shifting final request placement from the central manager to those groups. The manager performs only inter-group dispatch, while each selected group uses a dynamically selected primary worker to maintain group-local state and dispatch each invocation to the worker expected to provide the best response time within that group. Implemented on Apache OpenWhisk, FaaSHive improves throughput by up to 1.5 × over manager–worker FaaS scheduling baselines in our setup, while markedly reducing cold-start rates.