What Is an NPU in a Mini PC? TOPS, Copilot+ and Real Uses Explained

Every mini PC spec sheet in 2026 lists an NPU and a TOPS figure. This guide explains what the NPU actually does, why Microsoft requires 40 TOPS for Copilot+, and when the NPU matters — or does not — for your workloads.

By Stéphane Updated 10 min read · ~2,800 words

Affiliate disclosure: this article contains affiliate links. MiniPCDeals.net earns a commission on qualifying purchases at no extra cost to you.

In Short

The NPU is a dedicated silicon block on a modern processor designed for AI inference. It runs background AI tasks such as video-call background blur, real-time transcription and image upscaling at 5-10W instead of the 30-40W the CPU or GPU would consume. Microsoft requires at least 40 TOPS from an NPU to certify a device as a Copilot+ PC. For local LLM inference through Ollama or LM Studio, the NPU provides only marginal help: RAM capacity and memory bandwidth matter far more. In day-to-day use, the NPU handles Windows AI features silently in the background without affecting performance.

Key facts

  • An NPU (Neural Processing Unit) runs AI inference at 5-10W, versus 30-40W on a GPU and 30-65W on a CPU.
  • Microsoft requires 40 TOPS from an NPU to certify a device as Copilot+ PC.
  • AMD’s Ryzen AI 9 HX 370 and Ryzen AI Max+ 395 both ship the same XDNA 2 NPU rated at 50 TOPS.
  • The AMD Ryzen AI 400 Series (XDNA 2+) reaches 60 TOPS, announced at CES 2026.
  • Intel Core Ultra 200V has a 48 TOPS NPU; Qualcomm Snapdragon X Elite has 45 TOPS.
  • Intel N100 and N150 mini PCs have no dedicated NPU and cannot be Copilot+ certified.
  • The NPU has no measurable impact on gaming frame rates — that is purely GPU territory.
  • For local LLM inference, RAM capacity and memory bandwidth determine tokens per second, not NPU TOPS.
  • Six mini PCs in 2026 ship with a 50 TOPS XDNA 2 NPU: BOSGAME M5, GMKtec EVO-X3, Peladn HO5, Beelink SER9 Pro AI, ACEMAGIC Retro X5 and GMKtec EVO-X2.
  • TOPS is a synthetic INT8 metric — real-world AI performance also depends on software optimisation.
Copilot+ minimum
40 TOPS
Microsoft requirement
AMD HX 370
50 TOPS
XDNA 2 NPU
AMD AI 400
60 TOPS
XDNA 2+ (2026)
NPU power draw
5-10W
vs 30-40W on GPU

What is an NPU? The simple definition

An NPU (Neural Processing Unit) is a dedicated silicon block integrated into a modern processor that is designed specifically for one type of task: running AI inference efficiently. It operates in the background while the CPU and GPU handle their normal workloads.

The NPU is a specialised circuit built to execute the matrix multiplication operations that power AI models. Unlike a CPU, which excels at sequential logic, or a GPU, which excels at parallel graphics rendering, an NPU is architected specifically for lower-precision integer arithmetic — the type of math that AI inference requires. That specialisation makes NPUs dramatically more power-efficient than routing AI tasks through general-purpose hardware.

The key word is dedicated. Before NPUs existed, AI tasks on a PC had to run on the CPU, which is slow and power-hungry for this workload, or on the GPU, which is fast but draws 30-40W even for background tasks. An NPU performs the same AI work at 5-10W, which makes always-on AI features practical on compact devices like mini PCs without causing thermal issues.

NPUs are not new. Apple has shipped a Neural Engine since 2017. What changed in 2025-2026 is that AMD, Intel and Qualcomm all integrated powerful NPUs into their Windows PC chips, and Microsoft built a set of Windows 11 features that specifically require them.

NPU vs CPU vs GPU: what each one does

The CPU handles general logic and sequential tasks. The GPU handles graphics and parallel computation. The NPU handles AI inference efficiently in the background. On a modern mini PC, all three coexist on the same chip and share the same RAM pool.

Comparison of CPU, GPU and NPU on a mini PC
ComponentPrimary roleAI workload power draw
CPUGeneral logic, OS, applications, sequential processing30-65W
GPU (iGPU)Rendering, display output, gaming, parallel compute15-40W
NPUAI inference only — dedicated matrix math5-10W

The practical result: when Windows detects your face for auto-framing in a video call, transcribes speech in real time, or blurs your background, those tasks run on the NPU. The CPU stays free for your browser and your code, and the GPU stays free for rendering. AI features become nearly invisible in terms of system impact.

The MAC unit analogy

The core of an NPU is thousands of MAC (Multiply-Accumulate) units — circuits that perform the fundamental operation of AI inference (multiplying two numbers and adding them to a running total) extremely fast. The NPU also processes data at INT8 precision (8-bit integers) rather than the FP32 (32-bit floating point) that general-purpose hardware uses. INT8 requires 4× less memory and significantly less power, which is why NPU efficiency is so dramatically better for AI tasks.

What does TOPS mean?

TOPS stands for Tera Operations Per Second, or one trillion mathematical operations per second. It measures the raw throughput of an NPU at INT8 operations. Higher TOPS means faster AI models in theory, but TOPS is a synthetic benchmark metric, so real-world performance also depends on software optimisation and memory bandwidth.

The AMD Ryzen AI 9 HX 370’s NPU delivers 50 TOPS of INT8 throughput. That means it performs 50 trillion eight-bit multiply-accumulate operations every second — far more than any AI background task on a PC requires. The 50 TOPS figure exceeds Microsoft’s 40 TOPS Copilot+ requirement by 25%.

NPU TOPS comparison — 2026 mini PC processors

AMD Ryzen AI 400 Series (XDNA 2+)60 TOPS
Announced at CES 2026. Future mini PCs — not yet widely available.
AMD Ryzen AI 9 HX 370 / Max+ 395 (XDNA 2)50 TOPS
Current standard in 2026: BOSGAME M5, GMKtec EVO-X3, Peladn HO5, Beelink SER9 Pro AI.
Intel Core Ultra 200V (NPU 4)48 TOPS
Found in some Intel-based mini PCs. Slightly below AMD’s current offering.
Qualcomm Snapdragon X Elite45 TOPS
ARM-based, rare in traditional mini PCs. Common in Copilot+ laptops.
Microsoft Copilot+ minimum: 40 TOPS
Intel N100 / N150 (no dedicated NPU)0 TOPS
Budget mini PCs (GMKtec G3 Plus, GEEKOM Air12). Not Copilot+ certified. Fine for cloud AI use.
TOPS is a marketing metric — here is the honest context

TOPS measures peak INT8 throughput under synthetic conditions. Real-world AI performance depends on software optimisation, memory bandwidth and model architecture. AMD’s XDNA 2 benchmarks show 7-8% faster real-world AI inference than Intel’s 48 TOPS chip despite the narrower TOPS gap, because XDNA 2 has more on-chip memory and better spatial dataflow. Do not choose a mini PC based on TOPS alone — the platform’s software ecosystem matters as much as the number.

Copilot+ PC: what the 40 TOPS requirement actually unlocks

Microsoft requires 40+ TOPS NPU performance for a device to qualify as a Copilot+ PC and access a specific set of Windows 11 AI features. These features are exclusive to Copilot+ certified hardware and do not run on older machines, even high-end ones.

Copilot+ PC features and their NPU requirements
FeatureFunctionNPU requiredWorks without NPU
Windows RecallSearchable visual history of PC activityYes (40+ TOPS)No
Live Captions + TranslationReal-time transcription and translation of any audio, 44 languagesYes (40+ TOPS)No
Windows Studio EffectsBackground blur, auto-framing, eye contact, voice focus in callsYes (40+ TOPS)No
Cocreator (Paint)Local AI image generation from text promptsYes (40+ TOPS)No
Super Resolution (Photos)Local AI upscaling of photosYes (40+ TOPS)No
Microsoft CopilotAI assistant in Windows sidebarPartial (cloud fallback)Partially (cloud)
Adobe CC 2026Content-aware fill, auto-masking, generative expandRecommendedGPU fallback (slower)

If you use Windows video calls frequently — Zoom, Teams or Google Meet — and want background blur and noise cancellation without loading the CPU, an NPU-equipped mini PC makes a real difference. These effects ran on the CPU on older machines, consuming meaningful processing power during calls.

What the NPU actually does day-to-day in a mini PC

In normal daily use on a 2026 mini PC, the NPU operates entirely in the background. You do not interact with it directly. It silently handles AI workloads so the CPU and GPU remain available for the tasks you care about.

Video calls

This is where the NPU delivers the most tangible improvement for home office users. With a 50 TOPS NPU, Windows Studio Effects — background blur, auto-framing, eye contact correction — run entirely on the NPU at 5-10W. Without an NPU, these effects route through the CPU, consuming 15-20% of CPU capacity during video calls. On a 12-core HX 370 this is negligible; on a 4-core N100, it can noticeably slow other applications.

Transcription and translation

Live Captions with real-time translation transcribes any audio playing on your PC and translates it to another language in real time, entirely locally on the NPU. On a Copilot+ mini PC, you can transcribe a French YouTube video to English in real time with no internet connection and no API cost.

Windows Recall

Recall continuously captures screenshots of your screen, processes them through a local AI model to extract text and context, and makes everything you have done on your PC searchable by natural language. This runs on the NPU in the background. It is opt-in, and the data never leaves your device.

What the NPU does not do

Three common misconceptions

NPUs do not boost gaming FPS. Gaming performance is determined entirely by the GPU. An NPU has zero direct impact on frame rates in any current game.

NPUs do not run large language models efficiently. Running Llama 3.1 70B or Qwen3 32B requires the GPU and large amounts of RAM. The NPU can help with very small models, but for serious LLM inference, RAM bandwidth and GPU compute are the bottlenecks.

TOPS does not directly predict real-world AI speed. Software optimisation matters as much as hardware. An application not optimised for the AMD XDNA 2 architecture will not use the NPU at all.

NPU and local AI: does it actually help with LLMs?

For running local LLMs via Ollama or LM Studio, the NPU plays a minor supporting role. The primary hardware requirements are RAM for model size and memory bandwidth for tokens per second. The GPU is the main inference engine, not the NPU.

This is one of the most misunderstood aspects of AI PCs. Spec sheets prominently feature NPU TOPS, which leads many buyers to believe a higher TOPS rating means faster LLM performance. That is not accurate.

When you run a 7B or 70B language model through Ollama on a mini PC, the model weights load into RAM, and the GPU — in this case, the integrated Radeon 890M or 8060S — performs the matrix operations for each token generation. Memory bandwidth, meaning how fast the GPU reads from RAM, determines tokens per second. The NPU is not involved in this pipeline in any significant way.

The NPU does help with local AI in one case: for very small models specifically optimised for NPU execution, such as Microsoft’s Phi-3 Mini, the NPU runs inference at lower power than the GPU. But for the popular models used in Ollama — Llama, Mistral, Qwen, DeepSeek — the GPU path is faster and better supported.

What actually matters for local LLM inference

RAM capacity determines the maximum model size (16GB for 7B, 32GB for 14-32B, 128GB for 70B+).

Memory bandwidth determines tokens per second. The BOSGAME M5 and GMKtec EVO-X3, with 256 GB/s of bandwidth on Strix Halo, reach around 40-60 tokens per second on mid-size models. The Ryzen AI 9 HX 370 with 120 GB/s reaches around 35 tokens per second.

GPU compute units perform the actual inference calculations.

NPU has minimal role in current LLM toolchains. It matters for Copilot+ features, not for Ollama.

For a full breakdown of which mini PCs are best for running local AI models, see our best mini PC for local AI 2026 guide. To see how Strix Halo systems compare with the NVIDIA DGX Spark for local AI, read our DGX Spark 64GB vs 128GB vs Strix Halo comparison.

Which mini PCs have the best NPU in 2026?

Mini PCs with AMD’s XDNA 2 NPU at 50 TOPS offer the best combination of AI performance and availability in 2026. These include the BOSGAME M5 and GMKtec EVO-X3 (Ryzen AI Max+ 395), plus the Peladn HO5, Beelink SER9 Pro AI and ACEMAGIC Retro X5 (Ryzen AI 9 HX 370). Budget mini PCs with Intel N100 or N150 processors have no dedicated NPU.

Mini PCs with the best NPU in 2026
ModelProcessorMemoryNPUReview
BOSGAME M5Ryzen AI Max+ 395128GB50 TOPSReview
GMKtec EVO-X3Ryzen AI Max+ 395128GB50 TOPSReview
Peladn HO5Ryzen AI 9 HX 37032GB50 TOPSReview
Beelink SER9 Pro AIRyzen AI 9 HX 37032GB50 TOPSReview
ACEMAGIC Retro X5Ryzen AI 9 HX 37032-128GB50 TOPSReview
GMKtec EVO-X2Ryzen AI Max+ 395128GB50 TOPSReview
GMKtec G3 PlusIntel N15016GBNo NPUGuide

All six Ryzen AI mini PCs in the table share the same 50 TOPS NPU. There is no NPU performance difference between them. The choice comes down to price, form factor, memory configuration (16GB up to 128GB on Strix Halo), expandability and brand. For local AI work specifically, the 128GB Strix Halo machines — the BOSGAME M5 and GMKtec EVO-X3 — are the strongest options because memory capacity and bandwidth matter far more than NPU TOPS.

Do you actually need a high-TOPS NPU?

If you use Windows video calls regularly, want Copilot+ features, or work with AI-enhanced creative software such as Adobe CC 2026, a 40+ TOPS NPU is genuinely useful. For pure coding, web browsing, gaming or home server use, the NPU makes no difference at all.

Buy a mini PC with NPU if

  • You use video calls daily (Zoom, Teams, Meet) and want background blur and noise cancellation without loading the CPU
  • You want Windows Recall — searchable visual history of your PC activity
  • You want Live Captions with real-time translation running locally
  • You use Adobe CC 2026, which offloads AI operations to the NPU when available
  • You want a future-proof machine, as more apps add NPU support through 2026

Skip the NPU if

  • Your primary use is gaming, coding, home server, NAS or Plex — none of these use the NPU
  • You use a budget mini PC ($200-$400) mainly for browsing and Office — Copilot+ features are a bonus, not essential
  • You run local LLMs via Ollama — RAM and GPU compute matter far more than TOPS
  • You rely on cloud AI (ChatGPT, Claude) rather than local Windows AI features
The practical buying framework

If your budget is $800 or more and you are choosing between a Ryzen AI 9 HX 370 mini PC and an older Ryzen 7000-series machine at similar prices, choose the HX 370. The 50 TOPS NPU costs nothing extra at that price point and gives you Copilot+ certification, better power efficiency and a more future-proof platform. If your budget is under $500, do not stretch for the NPU — a well-specced N150 or Ryzen 4300U machine will serve most use cases perfectly. For a full roundup across every price tier, see our Best Mini PCs for Local AI in 2026 guide.

Best 50 TOPS mini PC

BOSGAME M5 — Ryzen AI Max+ 395 · 128GB · $2,999

50 TOPS XDNA 2 NPU, 128GB unified memory, 256 GB/s bandwidth. The cheapest 128GB Strix Halo mini PC in 2026.

Check price on Amazon

Frequently asked questions

What is an NPU in a mini PC?
An NPU (Neural Processing Unit) is a dedicated silicon block on a modern processor designed to run AI inference efficiently. It handles background AI tasks such as video-call background blur, real-time transcription and image upscaling at 5-10W, instead of the 30-40W the CPU or GPU would consume. On 2026 mini PCs, the NPU is typically AMD’s XDNA 2 architecture rated at 50 TOPS.
What does TOPS mean for an NPU?
TOPS stands for Tera Operations Per Second, or one trillion integer operations per second. It measures the raw INT8 throughput of an NPU. Microsoft requires at least 40 TOPS to certify a device as a Copilot+ PC. AMD’s Ryzen AI 9 HX 370 delivers 50 TOPS; the Ryzen AI 400 Series reaches 60 TOPS. TOPS is a synthetic metric, so real-world performance also depends on software optimisation and memory bandwidth.
Do I need a Copilot+ PC for a mini PC?
It depends on your workload. Copilot+ PC certification requires a 40+ TOPS NPU and unlocks specific Windows 11 features: Windows Recall, Live Captions with real-time translation, Windows Studio Effects and Cocreator. If you use these features, a 40+ TOPS mini PC is worth it. For web browsing, Office, Docker or coding, the NPU has no measurable impact.
Is an NPU the same as a GPU?
No. A GPU is optimised for graphics rendering and parallel floating-point computation; it can run AI workloads but at higher power consumption. An NPU uses low-precision integer math (INT8) designed specifically for AI inference. An NPU drawing 5-10W handles AI tasks that would cost 30-40W on a GPU. The NPU frees the GPU for rendering while processing AI features in the background.
Does an NPU help with running local AI models like Ollama?
Minimally. For LLMs like Llama, Mistral or Qwen running through Ollama or LM Studio, inference happens on the GPU (iGPU) using the unified RAM pool. Memory bandwidth and RAM capacity are the bottlenecks, not the NPU. This is particularly true on 128GB Strix Halo systems like the BOSGAME M5 or GMKtec EVO-X3, where the GPU and memory do the heavy lifting.
Does an NPU improve gaming performance?
No. Gaming performance is determined entirely by the GPU. The NPU has zero direct impact on frame rates in any current game. Marketing that links AI PC hardware with gaming usually refers to GPU-based features such as DLSS 4 (NVIDIA Tensor Cores) or AMD FSR, not the NPU.
Do budget mini PCs have an NPU?
Generally no. Mini PCs with Intel N-series processors (N100, N150) such as the GMKtec G3 Plus or GEEKOM Air12 have no dedicated NPU and are not Copilot+ certified. Older AMD Ryzen chips (4000, 5000 and 7000 series) also lack a Copilot+ capable NPU. The 40+ TOPS requirement is met by AMD Ryzen AI 300 and 400 series, Intel Core Ultra 200V series and Qualcomm Snapdragon X chips.
By Stéphane Published April 9, 2026 · Updated October 7, 2026

Methodology and sources

This guide is a specifications and analysis article, not a hands-on review. NPU TOPS figures are taken from AMD official press releases (CES 2026, MWC 2026), Intel product briefs, Qualcomm documentation and independent coverage by NotebookCheck and TechRadar. Copilot+ PC requirements are drawn from Microsoft’s official documentation. Power consumption estimates come from manufacturer technical documentation and independent hardware testing. Strix Halo performance and memory bandwidth figures are based on Level1Techs and Tom’s Hardware testing of the BOSGAME M5 and GMKtec EVO-X3. Every figure in this article has been cross-checked against at least two independent sources as of October 7, 2026. Where sources conflict, the more conservative figure is quoted. No hands-on NPU benchmarking was performed by MiniPCDeals.net for this article.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top