Line segment detection is a critical preprocessing step in embedded vision applications such as autonomous navigation, visual SLAM, and industrial inspection. Deep learning methods achieve high accuracy but require substantial resources, limiting their deployment on resource-constrained platforms. Classical algorithms are efficient but exhibit content-dependent latency. This paper presents a low-latency ASIC architecture for real-time line segment detection. The proposed design is based on the step-length algorithm and incorporates five ASIC-specific features: register-based line buffering with data reuse, multiplierless MCM-based filtering, 8-class angle quantization, a CAM-like associative memory for single-cycle matching, and an optimized duplicate removal mechanism. The architecture is fully pipelined and processes one pixel per clock cycle with deterministic latency. Synthesized in a 45nm CMOS process, the design achieves 325 FPS at VGA resolution and 48 FPS at Full HD, with 25.54 mW power consumption and 0.412 mm\textsuperscript{2} area. At 125 MHz, the throughput increases to 406 FPS at VGA resolution with 31.48 mW power consumption. Compared with a 90nm ASIC implementation based on the Line Hough Transform, the proposed design reduces power consumption by 49\% and delivers over 1.6 times higher frame rate. The architecture is well suited for edge-computing applications requiring real-time performance, low power, and minimal area.
Amir Hossein Jalilvand, P. Panahi, M. Najafi· 1 citation
Fuzzy logic systems are widely used for intelligent decision-making under uncertainty, offering interpretability and robustness across diverse applications. However, the growing demand for real-time edge intelligence has exposed the limitations of software-based fuzzy inference: unpredictable latency, excessive power consumption, and inefficient resource utilization. This has motivated extensive research into hardware acceleration, spanning platforms from custom analog circuits and digital ASICs to reconfigurable FPGAs and ultra-low-power microcontrollers. This survey presents the first comprehensive, platform-centric review of hardware fuzzy systems, systematically organizing the literature into three principal categories: FPGA-based implementations, ASIC and custom VLSI realizations, and embedded, IoT, and TinyML platforms. For each category, we analyze architectural organization, resource mapping strategies, implementation trade-offs, and key design challenges. Our cross-platform comparative analysis reveals that no single platform dominates across all metrics. FPGAs offer flexibility and rapid prototyping, ASICs deliver peak performance and energy efficiency, while embedded and TinyML systems balance low power and cost for edge deployment. Despite significant progress, critical research gaps persist: the absence of standardized benchmarks, limited scalability of rule bases, insufficient design automation, and limited support for online learning and emerging memory technologies. We outline future directions including in-memory fuzzy computing with memristive crossbars, integration with TinyML ecosystems, explainable hardware AI, and open-source design automation. This survey serves as a {reference for} researchers and practitioners working on hardware-enabled fuzzy intelligence.
Amir Hossein Jalilvand, P. Panahi, M. Najafi· 0 citations
Processing in memory (PIM) offers a compelling pathway to overcome the data movement bottleneck in modern AI and data-centric systems. This work introduces MITRA, a reconfigurable magnetic tunnel junction (MTJ)-based in-memory architecture that leverages stochastic computing (SC) to implement a broad class of transcendental and nonlinear functions directly within memory. By combining stochastic bit-stream processing with compact finite-state-machines (FSMs) embedded in MTJ-FinFET logic-in-memory structures, the proposed design achieves low-latency and power-efficient computation without external datapaths, unlike the binary counterparts. Circuit-level simulations in 14-nm FinFET technology verify correct state transitions, stable stochastic outputs, and predictable power profiles. Extensive evaluations demonstrate high accuracy even with short bit-streams. We further integrate the design into a neural-network classifier and develop an FSM-aware training strategy that compensates for approximation errors, achieving up to 96.9% classification accuracy on the UCI Optical Digit benchmark. Overall, MITRA provides a compact, reconfigurable platform for nonlinear processing in next-generation edge AI systems.
Farzad Razi, M. Moghadam, M. Najafi et al.· International Symposium on L...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.