A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section.
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Healthy fork ratio
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
git clone https://github.com/usamahz/cpu-performance-engineeringResumen de Awesome Lists
# CPU Performance Engineering  [](https://github.com/usamahz/cpu-performance-engineering/actions/workflows/links.yml) [](https://github.com/usamahz/cpu-performance-engineering/actions/workflows/quality.yml) [](#contents) [](misc/benchmarks/README.md) [](misc/mcp/README.md) [](LICENSE) [](https://github.com/usamahz/cpu-performance-engineering/stargazers) Making a program fast on a modern CPU means knowing what the core does with each instruction, where the time actually goes, and how to prove a change helped. This is the reading that gets you there, in the order that makes the next piece legible. **Scope.** x86 and Arm server parts, from one instruction through to serving a model on CPU. Not language runtimes, database internals, or anything above the socket. **Evidence.** Primary sources only: the paper, the specification, the vendor manual, the repository, or a report by the person who did the work. Any number, anywhere in this repository, carries all seven fields set out in [What earns a place](#what-earns-a-place), or it is not quoted. **Proof.** Fourteen of the sections end in a benchmark under [misc/benchmarks/](misc/benchmarks/README.md): C source, the build line, the machine, the raw numbers and the analysis, all committed. Run them yourself. **Plug it in.** The MCP server in [misc/mcp/](misc/mcp/README.md) turns this list into a CPU performance brain inside any AI client: the reading order, every rejected candidate with the rule it failed, every benchmark, and an index of the linked sources built on the reader's own machine. Point it at real work and the answers quote those sources and cite them. With [uv](https://docs.astral.sh/uv/) installed, one command adds it. **Claude Code** ```sh claude mcp add --scope user cpu-perf -- uvx cpu-perf ``` **Codex** ```sh codex mcp add cpu-perf -- uvx cpu-perf ``` Claude Desktop, Cursor and VS Code take a few lines of config, given in [Connect it](misc/mcp/README.md#connect-it). Section 1 is a path through the rest; read it top to bottom before using the numbered sections as a reference. ## Contents - [1. Start here](#1-start-here) - [2. One instruction, end to end](#2-one-instruction-end-to-end) - [Fetch and decode](#fetch-and-decode) - [Rename and issue](#rename-and-issue) - [Execute](#execute) - [Memory access and retire](#memory-access-and-retire) - [3. Microarchitecture](#3-microarchitecture) - [Limits of ILP and SMT](#limits-of-ilp-and-smt) - [Branch prediction and speculation](#branch-prediction-and-speculation) - [Vendor estimates and measured tables](#vendor-estimates-and-measured-tables) - [What the manuals leave out](#what-the-manuals-leave-out) - [4. Memory hierarchy](#4-memory-hierarchy) - [Cache geometry, replacement and misses in flight](#cache-geometry-replacement-and-misses-in-flight) - [TLBs, page walks and prefetchers](#tlbs-page-walks-and-prefetchers) - [Store buffers, ordering and cache-line contention](#store-buffers-ordering-and-cache-line-contention) - [Struct layout, software prefetch and page size](#struct-layout-software-prefetch-and-page-size) - [5. Measurement](#5-measurement) - [Method and the whole-system view](#method-and-the-whole-system-view) - [Counters, events and precise sampling](#counters-events-and-precise-sampling) - [CPU profilers and flame graphs](#cpu-profilers-and-flame-graphs) - [Microbenchmarks that lie](#microbenchmarks-that-lie) - [6. Models](#6-models) - [Roofline and the execution-cache-memory model](#roofline-and-the-execution-cache-memory-model) - [Top-down analysis](#top-down-analysis) - [Scaling laws](#scaling-laws) - [Queueing](#queueing) - [7. Single-thread optimisation](#7-single-thread-optimisation) - [Data layout and loop transforms](#data-layout-and-loop-transforms) - [SIMD instruction sets](#simd-instruction-sets) - [SIMD libraries and measured kernels](#simd-libraries-and-measured-kernels) - [Branchless code and bit manipulation](#branchless-code-and-bit-manipulation) - [8. Compilers and codegen](#8-compilers-and-codegen) - [Reading emitted code](#reading-emitted-code) - [Optimisation levels, inlining and link time](#optimisation-levels-inlining-and-link-time) - [Target flags and auto-vectorisation](#target-flags-and-auto-vectorisation) - [Profile-guided and post-link optimisation](#profile-guided-and-post-link-optimisation) - [9. Concurrency](#9-concurrency) - [Memory models and atomics](#memory-models-and-atomics) - [Locks, contention and allocators](#locks-contention-and-allocators) - [Lock-free structures and RCU](#lock-free-structures-and-rcu) - [Thread pools and work stealing](#thread-pools-and-work-stealing) - [10. NUMA and multi-socket](#10-numa-and-multi-socket) - [NUMA and Linux memory placement](#numa-and-linux-memory-placement) - [Topology and interconnects](#topology-and-interconnects) - [Migration, balancing and measured effects](#migration-balancing-and-measured-effects) - [11. OS and I/O](#11-os-and-io) - [Syscalls and asynchronous I/O](#syscalls-and-asynchronous-io) - [Scheduling, affinity and isolation](#scheduling-affinity-and-isolation) - [Interrupts and kernel bypass](#interrupts-and-kernel-bypass) - [Cache and bandwidth partitioning](#cache-and-bandwidth-partitioning) - [12. Tail latency and production systems](#12-tail-latency-and-production-systems) - [Measuring the tail](#measuring-the-tail) - [Where jitter comes from](#where-jitter-comes-from) - [Load generation and production workloads](#load-generation-and-production-workloads) - [Mechanical sympathy](#mechanical-sympathy) - [13. Inference on CPU](#13-inference-on-cpu) - [GEMM and BLAS](#gemm-and-blas) - [Runtimes](#runtimes) - [Quantization](#quantization) - [Matrix extensions](#matrix-extensions) - [Threading for inference](#threading-for-inference) - [When CPU beats GPU](#when-cpu-beats-gpu) - [14. Hardware generations](#14-hardware-generations) - [Intel Xeon](#intel-xeon) - [AMD EPYC](#amd-epyc) - [Arm Neoverse server parts](#arm-neoverse-server-parts) - [Independent measurement across vendors](#independent-measurement-across-vendors) - [15. Benchmarks](#15-benchmarks) - [Standard suites](#standard-suites) - [Microbenchmark suites](#microbenchmark-suites) - [Methodology and what suites miss](#methodology-and-what-suites-miss) - [16. Watchlist](#16-watchlist) - [ISA extensions without a shipped server part](#isa-extensions-without-a-shipped-server-part) - [Parts without a public measurement](#parts-without-a-public-measurement) - [Memory and interconnect](#memory-and-interconnect) - [Kernel paths and generated code](#kernel-paths-and-generated-code) - [What earns a place](#what-earns-a-place) ## 1. Start here Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names. 1. [Computer Architecture: A Quantitative Approach, 7th Edition](https://shop.elsevier.com/books/computer-architecture/hennessy/978-0-443-15406-5) - Its pipelining appendix and memory chapters define the hazard, speculation and cache vocabulary the list assumes. 2. [Optimizing software in C++](https://www.agner.org/optimize/optimizing_cpp.pdf) - Maps C++ onto pipeline mechanisms and shows why a loop-carried dependency chain, not instruction count, paces a loop. 3. [Intel Optimization Reference Manual](https://www.intel.com/content/www/us/en/content-details/671488/intel-64-and-ia-32-architectures-optimization-reference-manual-volume-1.html) - Its opening chapters show how a shipping x86 core implements the textbook pipeline, each rule tied to a mechanism. 4. [What Every Programmer Should Know About Memory](https://www.akkadia.org/drepper/cpumemory.pdf) - Measures the step in cost per access at each cache boundary and the gap a prefetcher hides. 5. [Memory Barriers: a Hardware View for Software Hackers](http://www.rdrop.com/users/paulmck/scalability/paper/whymb.2010.07.23a.pdf) - Explains why a second core makes loads and stores reorder and what a barrier drains. 6. [Systems Performance: Enterprise and the Cloud, 2nd Edition](https://www.brendangregg.com/systems-performance-2nd-edition-book.html) - Puts the method before the tools: what to measure, in what order, and how benchmarks mislead. 7. [Roofline: An Insightful Visual Performance Model for Multicore Architectures](https://cacm.acm.org/research/roofline-an-insightful-visual-performance-model-for-multicore-architectures/) - Places a loop from a byte count and a datasheet bandwidth alone, before any counter is read. 8. [A Top-Down Method for Performance Analysis and Counters Architecture](https://sites.google.com/site/analysismethods/yasin-pubs) - Defines the split of pipeline slots into front end, bad speculation, back end and retiring, the tree profilers report. 9. [Performance Analysis and Tuning on Modern CPUs](https://github.com/dendibakh/perf-book) - Walks from a noisy timing to counters to a named bottleneck, applying roofline and top-down to whole programs. 10. [What Has My Compiler Done for Me Lately? Unbolting the Compiler's Lid](https://www.youtube.com/watch?v=bSkpMdDe4g4) - Shows how to read emitted assembly against its source, so each mechanism is checked in a listing, not assumed. Work the exercises in [Performance Ninja](https://github.com/dendibakh/perf-ninja) alongside them; reading alon
Lo que la gente pregunta sobre cpu-performance-engineering
¿Qué es usamahz/cpu-performance-engineering?
+
usamahz/cpu-performance-engineering es awesome lists para el ecosistema de Claude AI. A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section. Tiene 663 estrellas en GitHub y su última actualización registrada es del 2026-10-05.
¿Cómo se instala cpu-performance-engineering?
+
Puedes instalar cpu-performance-engineering clonando el repositorio (https://github.com/usamahz/cpu-performance-engineering) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar usamahz/cpu-performance-engineering?
+
Nuestro agente de seguridad ha analizado usamahz/cpu-performance-engineering y le ha asignado un Trust Score de 100/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene usamahz/cpu-performance-engineering?
+
usamahz/cpu-performance-engineering es mantenido por usamahz. La última actividad registrada en GitHub es del 2026-10-05, con 2 issues abiertos.
¿Hay alternativas a cpu-performance-engineering?
+
Sí. En ClaudeWave puedes explorar awesome lists similares en /categories/awesome, ordenados por popularidad o actividad reciente.
Despliega cpu-performance-engineering en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/usamahz-cpu-performance-engineering)<a href="https://claudewave.com/repo/usamahz-cpu-performance-engineering"><img src="https://claudewave.com/api/badge/usamahz-cpu-performance-engineering" alt="Featured on ClaudeWave: usamahz/cpu-performance-engineering" width="320" height="64" /></a>Más Awesome Lists
A collection of MCP servers.
A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team at Anthropic PBC. A delectable showcase of top tier skills, ambidextrous agents, scintillating status lines, top notch developer tooling, and also we have plugins
AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes CLI, local MCP, catalog, plugins, and Workbench.
Your ultimate Go microservices framework for the cloud-native era.
A configuration framework that enhances Claude Code with specialized commands, cognitive personas, and development methodologies.