Skip to main content
Rolf Krause Research Group
Rolf Krause Research Group
Main navigation
Home
Applied Artificial Intelligence
EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models
Thu, Sep 25 2025
Research
Applied Artificial Intelligence
The deployment of language models is rapidly shifting from datacenters to edge devices such as laptops, smartphones, and embedded platforms, driven by the demand for interactive, low-latency, and privacy-preserving applications. In this context, Small Language Models (SLMs) have emerged as practical candidates, yet their inference reveals inefficiencies in conventional accelerators. While GPUs and NPUs process the GEMM-heavy prefill stage efficiently, they remain underutilized during the GEMV-dominated decoding phase, resulting in limited throughput and excessive energy consumption at the edge
Efficient AI Across Edge, Near-Edge, and Cloud
Thu, Sep 25 2025
Research
Applied Artificial Intelligence
Modern applications — smart cameras, self-driving cars, AR/VR headsets, on-device assistants — depend on increasingly large AI models. The catch is that the appetite of these models is growing far faster than the hardware meant to run them, especially at the edge. Closing that gap is not just a matter of building bigger chips; it requires rethinking where each part of a model executes across the devices a user actually has access to. Our work develops two complementary frameworks that address this directly. DONNA decides how to split a model across heterogeneous devices — CPUs, GPUs, and