Skip to content
Open access

LLMSGHD: Co-Designing Large Language Model Software and Hardware for Efficient Inference

Jul 2026 · ACM Transactions on Architecture and Code Optimization (TACO) · Vol 23, pp. 1 - 25 · 0 citations · 50 references

TL;DR

A hardware evaluation platform called LLMSGHD (Large Language Model Software-Guided Hardware Design), which focuses on operator-level optimized LLM inference workloads and integrates advanced software optimization techniques to offer more insightful analysis for hardware design.

Abstract

Over the past two years, the rapid evolution of Large Language Models (LLMs) and the relatively slower progress in hardware development have led to the emergence of many operator-level optimization techniques and theories. These software-based methods accelerate model inference by improving memory management and computational efficiency, and they have demonstrated empirical effectiveness. However, in current computer architecture research, hardware designers often focus on microarchitectural optimizations or the performance of individual operations. They tend to adopt existing software optimizations passively, without actively leveraging them as design principles. Constrained by this limited perspective, researchers may overlook opportunities to systematically integrate insights from software into hardware design, potentially hindering more efficient and flexible architectural innovations. To combine operator-level optimization methods with hardware design, we introduce a hardware evaluation platform called LLMSGHD (Large Language Model Software-Guided Hardware Design), which focuses on operator-level optimized LLM inference workloads. LLMSGHD simulates hardware inference behavior while aiming for broad applicability and high efficiency. LLMSGHD integrates advanced software optimization techniques to offer more insightful analysis for hardware design, particularly regarding computational density variations in inference and their interaction with software-level optimizations. Based on LLMSGHD, we develope a heterogeneous LLM inference platform targeting high throughput and low cost.

Read PDF

Similar papers

Preprint Aug 2026

Effect of Abstractions and Prompting Strategies on LLM-Guided High-Performance Optimizations

It is demonstrated that LLMs provided with specific optimization goals achieve better measured performance and validity rates when generating C code compared to creating computation pipelines and optimization schedules with established frameworks, suggesting that future development should explore alternative approaches...

Jiří Klepl, Matyáš Brabec, Martin Kruliš · 0 citations
Jul 2026

Profiling Lightweight Large Language Models

These results show that selecting lightweight LLMs by size, FLOPs, latency, or accuracy alone can select the wrong deployment candidate; PTME profiling exposes configurations that preserve useful accuracy at lower physical cost.

Tomohiro Harada, Enrique Alba, Gabriel Luque · 0 citations

Performance Evaluation of LLM Inference Engines

This thesis presents a performance evaluation of three widely used open-source inference engines: vLLM, SGLang, and llama.cpp, and summarizes the experimental findings into selection recommendations for practical deployment scenarios, providing a reference for developers and researchers in choosing an appropriate infer...

Unknown authors · 0 citations
#artificial intelligence Preprint Sep 2026

PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical...

Da Zhao, K. Sankaralingam, Christos Kozyrakis et al. · 0 citations
Preprint Aug 2026

LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization

LLM4LLM is introduced, a deployment-aware closed-loop optimization framework that starts from a target inference script, extracts phase-aware optimization tasks, searches with an experience-guided episodic agent, and accepts patches through in-model validation.

Hui Zeng, Pengfei Yang, Yanxin Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.