FlexXNOR-OD: A Channel-Grouped XNOR-Based Variable-Precision Accelerator for Real-Time Edge Object Detection
Abstract
Object detection at the edge requires a difficult balance among detection accuracy, deterministic latency, memory bandwidth, and energy consumption. Existing binarized accelerators replace multipliers with XNOR and population-count logic, but many designs use a fixed binary datapath or select precision only at the layer level, limiting their ability to protect accuracy-sensitive channels while fully exploiting low-bit parallelism. This paper proposes FlexXNOR-OD, a new FPGA-oriented accelerator for real-time object detection that combines channel-group precision assignment with a lane-fusible XNOR processing element. The architecture supports 1-, 2-, 4-, and 8-bit weight/activation groups using four independently clock-gated XNOR-popcount slices per processing element. The slices operate independently in binary mode and are fused through signed bit-plane recombination for higher-precision groups. A 64-PE output-stationary array, dual-bank feature and weight buffers, a precision-and-tile scheduler, and an integrated batch-normalization, RPReLU, and requantization pipeline minimize data movement and control overhead. An analytical model for a 200-MHz design point predicts a peak rate of 4096 binary dot-product bit pairs per cycle, equivalent to 0.819 Tbinary-MAC/s or 1.638 TOPS when a multiply and accumulation are counted separately. For an illustrative channel-group precision map with an average weight width of 2.12 bits, parameter storage is reduced by 3.77× relative to uniform INT8 and 15.1× relative to FP32. The proposed design therefore offers a practical path toward precision-scalable, multiplier-light object detection on resource-constrained edge platforms. A complete RTL synthesis and dataset-level evaluation protocol is specified to support reproducible implementation