Skip to main content
ClaudeWave
usamahz avatar
usamahz

cpu-performance-engineering

View on GitHub

A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section.

Awesome ListsOfficial Registry663 stars57 forks● PythonMITUpdated today
ClaudeWave Trust Score
100/100
✓ Verified
Passed
  • ✓Open-source license (MIT)
  • ✓Actively maintained (<30d)
  • ✓Healthy fork ratio
  • ✓Clear description
  • ✓Topics declared
  • ✓Documented (README)
Last scanned: 10/5/2026
Use this list
Method: Clone
Terminal
git clone https://github.com/usamahz/cpu-performance-engineering
1. Browse the curated list on GitHub or clone it locally.
2. Star it to keep new additions on your radar.
Use cases

Awesome Lists overview

# CPU Performance Engineering

![](misc/banner-pinnacle-ridge.avif)

[![Links](https://github.com/usamahz/cpu-performance-engineering/actions/workflows/links.yml/badge.svg)](https://github.com/usamahz/cpu-performance-engineering/actions/workflows/links.yml) [![Quality](https://github.com/usamahz/cpu-performance-engineering/actions/workflows/quality.yml/badge.svg)](https://github.com/usamahz/cpu-performance-engineering/actions/workflows/quality.yml) [![Entries](https://img.shields.io/badge/entries-304-1f6feb)](#contents) [![Benchmarks](https://img.shields.io/badge/benchmarks-14%20runnable-1f6feb)](misc/benchmarks/README.md) [![MCP](https://img.shields.io/badge/MCP-server-1f6feb)](misc/mcp/README.md) [![License](https://img.shields.io/badge/license-MIT-lightgrey)](LICENSE) [![Stars](https://img.shields.io/github/stars/usamahz/cpu-performance-engineering?style=flat&color=555)](https://github.com/usamahz/cpu-performance-engineering/stargazers)

Making a program fast on a modern CPU means knowing what the core does with
each instruction, where the time actually goes, and how to prove a change
helped. This is the reading that gets you there, in the order that makes the
next piece legible.

**Scope.** x86 and Arm server parts, from one instruction through to serving
a model on CPU. Not language runtimes, database internals, or anything above
the socket.

**Evidence.** Primary sources only: the paper, the specification, the vendor
manual, the repository, or a report by the person who did the work. Any
number, anywhere in this repository, carries all seven fields set out in
[What earns a place](#what-earns-a-place), or it is not quoted.

**Proof.** Fourteen of the sections end in a benchmark under
[misc/benchmarks/](misc/benchmarks/README.md): C source, the build line, the
machine, the raw numbers and the analysis, all committed. Run them yourself.

**Plug it in.** The MCP server in [misc/mcp/](misc/mcp/README.md) turns
this list into a CPU performance brain inside any AI client: the reading
order, every rejected candidate with the rule it failed, every benchmark,
and an index of the linked sources built on the reader's own machine. Point
it at real work and the answers quote those sources and cite them. With
[uv](https://docs.astral.sh/uv/) installed, one command adds it.

**Claude Code**

```sh
claude mcp add --scope user cpu-perf -- uvx cpu-perf
```

**Codex**

```sh
codex mcp add cpu-perf -- uvx cpu-perf
```

Claude Desktop, Cursor and VS Code take a few lines of config, given in
[Connect it](misc/mcp/README.md#connect-it).

Section 1 is a path through the rest; read it top to bottom before using the
numbered sections as a reference.

## Contents

- [1. Start here](#1-start-here)
- [2. One instruction, end to end](#2-one-instruction-end-to-end)
  - [Fetch and decode](#fetch-and-decode)
  - [Rename and issue](#rename-and-issue)
  - [Execute](#execute)
  - [Memory access and retire](#memory-access-and-retire)
- [3. Microarchitecture](#3-microarchitecture)
  - [Limits of ILP and SMT](#limits-of-ilp-and-smt)
  - [Branch prediction and speculation](#branch-prediction-and-speculation)
  - [Vendor estimates and measured tables](#vendor-estimates-and-measured-tables)
  - [What the manuals leave out](#what-the-manuals-leave-out)
- [4. Memory hierarchy](#4-memory-hierarchy)
  - [Cache geometry, replacement and misses in flight](#cache-geometry-replacement-and-misses-in-flight)
  - [TLBs, page walks and prefetchers](#tlbs-page-walks-and-prefetchers)
  - [Store buffers, ordering and cache-line contention](#store-buffers-ordering-and-cache-line-contention)
  - [Struct layout, software prefetch and page size](#struct-layout-software-prefetch-and-page-size)
- [5. Measurement](#5-measurement)
  - [Method and the whole-system view](#method-and-the-whole-system-view)
  - [Counters, events and precise sampling](#counters-events-and-precise-sampling)
  - [CPU profilers and flame graphs](#cpu-profilers-and-flame-graphs)
  - [Microbenchmarks that lie](#microbenchmarks-that-lie)
- [6. Models](#6-models)
  - [Roofline and the execution-cache-memory model](#roofline-and-the-execution-cache-memory-model)
  - [Top-down analysis](#top-down-analysis)
  - [Scaling laws](#scaling-laws)
  - [Queueing](#queueing)
- [7. Single-thread optimisation](#7-single-thread-optimisation)
  - [Data layout and loop transforms](#data-layout-and-loop-transforms)
  - [SIMD instruction sets](#simd-instruction-sets)
  - [SIMD libraries and measured kernels](#simd-libraries-and-measured-kernels)
  - [Branchless code and bit manipulation](#branchless-code-and-bit-manipulation)
- [8. Compilers and codegen](#8-compilers-and-codegen)
  - [Reading emitted code](#reading-emitted-code)
  - [Optimisation levels, inlining and link time](#optimisation-levels-inlining-and-link-time)
  - [Target flags and auto-vectorisation](#target-flags-and-auto-vectorisation)
  - [Profile-guided and post-link optimisation](#profile-guided-and-post-link-optimisation)
- [9. Concurrency](#9-concurrency)
  - [Memory models and atomics](#memory-models-and-atomics)
  - [Locks, contention and allocators](#locks-contention-and-allocators)
  - [Lock-free structures and RCU](#lock-free-structures-and-rcu)
  - [Thread pools and work stealing](#thread-pools-and-work-stealing)
- [10. NUMA and multi-socket](#10-numa-and-multi-socket)
  - [NUMA and Linux memory placement](#numa-and-linux-memory-placement)
  - [Topology and interconnects](#topology-and-interconnects)
  - [Migration, balancing and measured effects](#migration-balancing-and-measured-effects)
- [11. OS and I/O](#11-os-and-io)
  - [Syscalls and asynchronous I/O](#syscalls-and-asynchronous-io)
  - [Scheduling, affinity and isolation](#scheduling-affinity-and-isolation)
  - [Interrupts and kernel bypass](#interrupts-and-kernel-bypass)
  - [Cache and bandwidth partitioning](#cache-and-bandwidth-partitioning)
- [12. Tail latency and production systems](#12-tail-latency-and-production-systems)
  - [Measuring the tail](#measuring-the-tail)
  - [Where jitter comes from](#where-jitter-comes-from)
  - [Load generation and production workloads](#load-generation-and-production-workloads)
  - [Mechanical sympathy](#mechanical-sympathy)
- [13. Inference on CPU](#13-inference-on-cpu)
  - [GEMM and BLAS](#gemm-and-blas)
  - [Runtimes](#runtimes)
  - [Quantization](#quantization)
  - [Matrix extensions](#matrix-extensions)
  - [Threading for inference](#threading-for-inference)
  - [When CPU beats GPU](#when-cpu-beats-gpu)
- [14. Hardware generations](#14-hardware-generations)
  - [Intel Xeon](#intel-xeon)
  - [AMD EPYC](#amd-epyc)
  - [Arm Neoverse server parts](#arm-neoverse-server-parts)
  - [Independent measurement across vendors](#independent-measurement-across-vendors)
- [15. Benchmarks](#15-benchmarks)
  - [Standard suites](#standard-suites)
  - [Microbenchmark suites](#microbenchmark-suites)
  - [Methodology and what suites miss](#methodology-and-what-suites-miss)
- [16. Watchlist](#16-watchlist)
  - [ISA extensions without a shipped server part](#isa-extensions-without-a-shipped-server-part)
  - [Parts without a public measurement](#parts-without-a-public-measurement)
  - [Memory and interconnect](#memory-and-interconnect)
  - [Kernel paths and generated code](#kernel-paths-and-generated-code)
- [What earns a place](#what-earns-a-place)

## 1. Start here

Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.

1. [Computer Architecture: A Quantitative Approach, 7th Edition](https://shop.elsevier.com/books/computer-architecture/hennessy/978-0-443-15406-5) - Its pipelining appendix and memory chapters define the hazard, speculation and cache vocabulary the list assumes.
2. [Optimizing software in C++](https://www.agner.org/optimize/optimizing_cpp.pdf) - Maps C++ onto pipeline mechanisms and shows why a loop-carried dependency chain, not instruction count, paces a loop.
3. [Intel Optimization Reference Manual](https://www.intel.com/content/www/us/en/content-details/671488/intel-64-and-ia-32-architectures-optimization-reference-manual-volume-1.html) - Its opening chapters show how a shipping x86 core implements the textbook pipeline, each rule tied to a mechanism.
4. [What Every Programmer Should Know About Memory](https://www.akkadia.org/drepper/cpumemory.pdf) - Measures the step in cost per access at each cache boundary and the gap a prefetcher hides.
5. [Memory Barriers: a Hardware View for Software Hackers](http://www.rdrop.com/users/paulmck/scalability/paper/whymb.2010.07.23a.pdf) - Explains why a second core makes loads and stores reorder and what a barrier drains.
6. [Systems Performance: Enterprise and the Cloud, 2nd Edition](https://www.brendangregg.com/systems-performance-2nd-edition-book.html) - Puts the method before the tools: what to measure, in what order, and how benchmarks mislead.
7. [Roofline: An Insightful Visual Performance Model for Multicore Architectures](https://cacm.acm.org/research/roofline-an-insightful-visual-performance-model-for-multicore-architectures/) - Places a loop from a byte count and a datasheet bandwidth alone, before any counter is read.
8. [A Top-Down Method for Performance Analysis and Counters Architecture](https://sites.google.com/site/analysismethods/yasin-pubs) - Defines the split of pipeline slots into front end, bad speculation, back end and retiring, the tree profilers report.
9. [Performance Analysis and Tuning on Modern CPUs](https://github.com/dendibakh/perf-book) - Walks from a noisy timing to counters to a named bottleneck, applying roofline and top-down to whole programs.
10. [What Has My Compiler Done for Me Lately? Unbolting the Compiler's Lid](https://www.youtube.com/watch?v=bSkpMdDe4g4) - Shows how to read emitted assembly against its source, so each mechanism is checked in a listing, not assumed.

Work the exercises in [Performance Ninja](https://github.com/dendibakh/perf-ninja) alongside them; reading alon
armawesome-listbenchmarkscomputer-architecturecpuedge-aiinferencelow-latencyoptimizationperformanceperformance-engineeringsimdsystems-programmingx86

What people ask about cpu-performance-engineering

What is usamahz/cpu-performance-engineering?

+

usamahz/cpu-performance-engineering is awesome lists for the Claude AI ecosystem. A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section. It has 663 GitHub stars and its last recorded update is dated 2026-10-05.

How do I install cpu-performance-engineering?

+

You can install cpu-performance-engineering by cloning the repository (https://github.com/usamahz/cpu-performance-engineering) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.

Is usamahz/cpu-performance-engineering safe to use?

+

Our security agent has analyzed usamahz/cpu-performance-engineering and assigned a Trust Score of 100/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.

Who maintains usamahz/cpu-performance-engineering?

+

usamahz/cpu-performance-engineering is maintained by usamahz. The last recorded GitHub activity is dated 2026-10-05, with 2 open issues.

Are there alternatives to cpu-performance-engineering?

+

Yes. On ClaudeWave you can browse similar awesome lists at /categories/awesome, sorted by popularity or recent activity.

Deploy cpu-performance-engineering to your cloud

Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.

Maintain this repo? Add a badge to your README

Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.

Featured on ClaudeWave: usamahz/cpu-performance-engineering
[![Featured on ClaudeWave](https://claudewave.com/api/badge/usamahz-cpu-performance-engineering)](https://claudewave.com/repo/usamahz-cpu-performance-engineering)
<a href="https://claudewave.com/repo/usamahz-cpu-performance-engineering"><img src="https://claudewave.com/api/badge/usamahz-cpu-performance-engineering" alt="Featured on ClaudeWave: usamahz/cpu-performance-engineering" width="320" height="64" /></a>