Ana içeriğe atla

HARDWARE ORIENTED MODELING OF LLM ACCELERATOR PIPELINE

tarihinde Adsız tarafından gönderildi
MSc Thesis📅 22.07.2026 — 13:30
👤 Speaker:
ARDA CEKIC
🎓 Supervisor(s):
ASSOC.PROF.DR.SEYDA ERTEKIN
📍 Location:
METU DTX
⏲ Duration:
90 min.
📝 Abstract:

Large Language Models (LLMs) achieve strong performance in many natural language processing tasks, but their high computational and memory requirements make hardware deployment challenging. Quantization reduces these costs by using lowprecision weights and activations, yet practical accelerator design also requires modeling integer arithmetic, accumulator precision, re-scaling, quantization-domain transitions, and data movement. This thesis presents a Python-based hardware-oriented modeling framework for quantized LLM inference based on SmoothQuant-style posttraining quantization. The framework models key decoder stages, including Normalization layers, projection layers, attention matrix multiplications, MLP operations, element-wise operations, INT8 operands, INT32 accumulation, and output re-scaling. It is integrated into a Neural Compressor-based evaluation flow and evaluated on LLaMA-2-7B and OPT-1.3B using LAMBADA, Winogrande, PIQA, and WikiText. The results show only minor degradation compared with the baseline quantized flow, while sparsity in the integer domain is also analyzed for future sparse-aware hardware design.

Time - Location
2026-07-22 13:30:00