WPR/WPA in Practice — A Primer on Investigating "the Whole PC Is Slow" Across the Entire System

· Updated: · · Windows, Performance, WPR, WPA, ETW, Performance Investigation, Bug Investigation, Windows Development

Revision history (1 updates, last updated Sep 8, 2026)

A log of the changes made to this article. Where a pre-update version was archived, it stays readable at a permanent DOI link.

Retranslated as a full translation of the current Japanese original. The previous English version was an abridgement that dropped subsections, tables, diagrams, and paragraphs; all of them have been restored to match the Japanese article, and the knowledge map section has been added where the Japanese article has one. The technical claims are the same as in the Japanese version. Read the version before this update (DOI: 10.5281/zenodo.22170889)
First published
Cite this article(DOI: 10.5281/zenodo.22170888)

This article is archived on Zenodo. Below are both the DOI that always resolves to the latest version and the DOI pinned to the version you are reading.

Go Komura (2026). WPR/WPA in Practice — A Primer on Investigating "the Whole PC Is Slow" Across the Entire System. KomuraSoft LLC. https://doi.org/10.5281/zenodo.22170888 https://comcomponent.com/en/blog/wpr-wpa-system-performance-analysis/

DOI (latest version)
10.5281/zenodo.22170888
DOI (this version)
10.5281/zenodo.22652536

“The whole PC got slow after we installed an app, but Task Manager shows headroom in both CPU and memory.” “Just one machine takes 3 minutes to boot.” What makes these consultations hard is that you do not even know which process to look at.

Even when app A is the slow one, the cause may be an antivirus scan, heavy writes by another service, or a chain of locks spanning several processes. What you need is to record the activity of the whole OS on a single timeline and follow where the time went.

The tools for that are Windows Performance Recorder (WPR) and Windows Performance Analyzer (WPA). WPR captures the recording using ETW (Event Tracing for Windows), and WPA reads that recording in graphs and tables. You examine, along the timeline, which processes and stacks used CPU, what each thread was waiting for, and which process issued disk I/O to which file.

The question this article answers is “how do you record system-wide slowness, and which of CPU, wait time, and I/O do you examine?” Written for IT staff at small and medium businesses and for Windows app developers, it covers everything from the practicalities of capture to analysis, based on primary sources as of August 2026. Read it with the difference between “CPU is high” and “CPU is low but it is still slow” as the guiding axis.

How this differs from Process Monitor, which examines file and registry access, and PerfView, which follows a .NET app’s CPU and GC, is covered in Chapter 2.

1. The Bottom Line First

The basic flow is “capture with WPR → narrow to the problem interval in WPA → isolate CPU, waits, or I/O.” Even a problem that a single process cannot explain can be investigated by following every process and the kernel on the same timeline.12

Separate Where You Capture From Where You Read

The capture tool wpr.exe ships with Windows 8.1 and later, so no additional install is needed. The GUI edition, WPRUI, and the analysis tool WPA are included in the Windows ADK. Because you can split the work as “only capture with wpr.exe in the customer environment; read it in WPA on your own machine,” you can capture even on a server where no software can be added.12

For an issue you can reproduce, the basic three steps are, with administrator privileges, wpr -start GeneralProfile -filemode → reproduce the issue → wpr -stop C:\temp\trace.etl.3 Prepare the destination folder, get approval for the capture, and decide how the ETL will be handled before you run it. Waiting for an issue and problems during boot call for different capture methods, so see Chapter 3 and Chapter 8.

The First Graph to Open in WPA

First narrow to the interval of the issue, then pick the direction of investigation from the following table.45

State in that interval What to look at first What to confirm
CPU is high Chapter 5: CPU Usage (Sampled) Which process, stack, and function used the CPU
Overall CPU is low, but one core or one thread is pinned Chapter 5: CPU Usage (Sampled) Whether a CPU bottleneck is buried under the overall utilization
Slow even though neither the CPU nor any particular core is pinned Chapter 6: CPU Usage (Precise) Where it waited, how long it waited, and who released the wait
Disk or file access is suspected Chapter 7: Disk Usage / File I/O Queue time versus device service time, and which process issued I/O to which file

Sampled shows the breakdown of CPU time from samples taken about every millisecond.6 Precise follows waits from a complete record of context switches. Tracing the party that released a wait, using Waits → ReadyingProcess → ReadyThreadStack as the clues, is the technique this article most wants to convey.47

How to Read This Article by Goal

If you are in charge of capturing, start with Chapter 2 and Chapter 3; if you are reading an ETL that has already been captured, start with Chapter 4. Reading stacks by function name requires symbol configuration, and for your own app you add the path to its PDBs.8 When you investigate .NET JIT code, CLR events at capture time are also required, so check the .NET notes in Chapter 4 before capturing.

The overall investigation workflow and ETL handling are summarized in Chapter 9. An ETL contains internal information such as process names and file paths, so keep the capture to the minimum required and decide the conditions for handing it outside the company in advance.

In the diagram a solid line marks a relation that always holds and a dashed line marks a conditional one (the conditions are given per relation on the detail page). The full list of relations (16 in total, with evidence and certainty) and the definitions of the main concepts are collected on the knowledge map detail page (in Japanese). Data: JSON-LD / Turtle

2. Where the Tools Sit — WPR Captures, WPA Reads

Windows Performance Toolkit (WPT) is the set of performance investigation tools included in the Windows ADK (Windows Assessment and Deployment Kit), and its core is the pair WPR and WPA.2 Their roles are clearly separated.

WPR Captures, WPA Analyzes

  • WPR (Windows Performance Recorder) = capture. It bundles ETW providers into units called “profiles,” starts and stops recording, and produces an ETL file. The command-line edition, wpr.exe, ships with Windows 8.1 and later, with no additional install. The GUI edition (WPRUI.exe) is included in the ADK.1
  • WPA (Windows Performance Analyzer) = analysis. It opens an ETL file and analyzes it in graphs and tables. Installing the ADK is required.2

In other words, if all you do is capture from the command line, you do not need to install any additional software in the customer environment. Capture with the OS-standard wpr.exe, take the ETL file back, and read it in WPA on your own PC — the same split as “capture with pktmon, read in Wireshark” in packet capture.

The split between capturing with WPR and reading with WPAIn the customer environment, record with the OS-standard wpr.exe to produce an ETL file, take it back, and analyze it in WPA installed via the ADK on your own PCYour own PC (WPA installed via the ADK)Customer environment (no additional install)Take backGraph and table analysis in WPAETL filewpr -start → reproduce the issue → wpr -stop

Choosing Between Procmon, PerfView, and WPR/WPA

Let us also sort out first how this differs from similar tools.

  Process Monitor PerfView WPR + WPA
The question it answers Which process did what to which path, and what the result was What a .NET app’s CPU, GC, and allocations are doing Where, across the whole OS, the time went
Coverage Operation log of file, registry, and process-start activity Mainly managed code The whole system: CPU, waits, disk, file I/O, power, and more
Symptoms it suits Settings not being read, ACCESS DENIED Slowness or memory of your own .NET app alone The whole PC is slow, CPU is idle but it is slow, culprit process unknown
Article Procmon Practical Guide PerfView Practical Introduction This article

If Procmon is the operation log of “what it did” and PerfView is “what happened inside .NET,” WPA is the tool that audits “where the time went” across every process.

When You Want to Record Your Own App’s Events Too

How ETW itself works and how to instrument your own app with ETW are covered in “An Introduction to Windows Event Log and ETW.” If your app emits ETW events, its milestones are recorded in the same trace, which makes correlation far easier.

Note, however, that WPR records only the events of providers enabled by the recording profile you selected. GeneralProfile does not include your own provider. To record them together, prepare a custom recording profile (.wprp) that enables your provider and combine it as in wpr -start GeneralProfile -start MyApp.wprp!MyAppProfile, specifying the profile name inside the .wprp file after !3.

3. Capture in Practice (WPR) — start, Reproduce, stop

What to Decide Before Running It

In a production environment, do not assume that even a short capture is unconditionally safe. ETW is lightweight, but recording a large volume of events with stacks consumes a certain amount of CPU and memory. Choose a time with little impact on business, and put it through the same approval process as any ordinary change.

If you can reproduce the issue, start just before, stop just after, and keep it within a few minutes. Prepare the destination folder, and decide in advance who receives the ETL, how long it is retained, and how it is deleted. Cautions about its contents are in Chapter 9.

The following is an example of capturing an issue you can reproduce on the spot in File mode. For an issue whose timing is unknown, go to Memory mode in Section 3.1; for problems during boot or logon, go to Chapter 8.

A Normal Capture Is “start → Reproduce → stop”

Run these in a Command Prompt opened as administrator. -profiles lists the available profiles beforehand, and -status checks the state while capturing. The -cancel at the end is used only to abort midway and discard; it is not part of the normal save procedure.

:: List the built-in profiles you can use
wpr -profiles

:: 1. Start the capture (general-purpose profile, file mode)
wpr -start GeneralProfile -filemode

:: 2. Reproduce the issue (check the capture state with wpr -status)

:: 3. Stop and save (you can attach a description of the problem).
::    Create the destination folder in advance (without it, -stop fails to save)
mkdir C:\temp 2>nul
wpr -stop C:\temp\slow-pc.etl "Reproduced the whole PC becoming slow while App X starts"

:: To abort midway, discard without saving
wpr -cancel

Choose What to Record With a Profile

What you pass to -start is a profile, a bundle of the ETW providers that investigation needs.3 Remembering only the frequently used ones is enough.9

Profile What it records When to use it
GeneralProfile A general-purpose set: CPU samples, context switches, disk I/O, and more Start here. The first move when you do not yet know what is wrong
CPU Details of CPU usage When you already know the CPU is burning
DiskIO Disk I/O activity When the disk is suspected
FileIO File I/O activity When you need to follow which file is being accessed

You can specify several profiles at once by repeating -start (for example, wpr -start GeneralProfile -start FileIO -filemode).3

WPR capture flow and how to choose a modeAn issue you can reproduce locally is captured briefly and reliably in file mode; an issue whose timing is unknown is waited for in the default memory-mode ring buffer. An issue during boot or logon uses a boot trace. In every case the start, reproduce, stop procedure is the sameReproducible locallyTiming unknownDuring boot or logonWhen does the issue occur?Capture briefly and reliably with -filemodeWait for it in Memory mode (default, ring buffer) (Section 3.1)Boot trace (Chapter 8)wpr -start → reproduce or wait for the issue → wpr -stop

3.1. Memory Mode and File Mode — Can You Reproduce It, or Do You Wait for It?

WPR has two recording destination modes, and the default is Memory mode (an in-memory circular buffer).

Memory mode is for waiting for an issue to occur. Because it is a ring buffer that overwrites the oldest events first, it suits leaving the capture running while you wait for an issue whose timing is unknown, and stopping once it happens.

File mode is for issues you can reproduce reliably in a short time. Adding -filemode switches to File mode, and everything is recorded to a continuous file. Nothing is lost to overwriting, but in exchange the only ceiling is free disk space, and the file grows without bound.10

How Memory mode and File mode recordMemory mode records to an in-memory circular buffer where the oldest events are overwritten and only the most recent remain, so it suits waiting for an issue. File mode keeps everything in a file, but the only ceiling is free disk space, so it suits a short, reliable reproductionStream of ETW eventsMemory mode: circular buffer (oldest overwritten first → only the most recent remain)File mode: everything kept in a file (ceiling is free disk space)Suits waiting for an issue whose timing is unknownSuits an issue you can reproduce reliably in a short time

The rule of thumb for choosing is as follows.

  • Reproducible on the spot → File mode. Start just before the reproduction, stop just after, and keep the capture within a few minutes
  • Timing unknown → Wait in Memory mode (the default). As soon as it occurs, run wpr -stop
  • Even a few minutes of GeneralProfile can produce an ETL in the hundreds of MB to GB range. A file that is too large may not be analyzable in WPA, so “the longer the capture, the better” is counterproductive1011

Capturing From the GUI

To capture from the GUI, start WPRUI, choose a profile and Logging mode, and click Start and then Save. The details of the procedure are in the official How-to topics.11 When you ask a contact at the customer site to capture, you can also write a procedure document centered on the start, reproduce, and stop steps, with the prerequisites and the abort operation set out separately.

4. The Basics of Reading WPA — Graphs, the Golden Rule of Tables, and Narrowing the Time Range

When you open the captured ETL in WPA, the Graph Explorer on the left lists graph thumbnails in categories such as System Activity, Computation, Storage, and Memory.12 Drag the graph you want onto the Analysis tab on the right, and the graph appears at the top with a table below it.

The three things to master are column ordering, narrowing the time range, and symbol configuration. Rather than relearning the operations for each graph, start from this shared way of reading.

4.1. Column Order Decides How Data Is Aggregated

A WPA table has two vertical bars, one gold and one blue: the columns to the left of the gold bar hierarchize (group) the data in that order, and the columns to the right of the blue bar are aggregate values.13

Ordering them “Process → Stack” gives a stack aggregation per process; “Stack → Process” gives an aggregation across every process that uses the same stack. Dragging columns to reorder them is itself the analysis operation. Once you understand this one point, every WPA table reads the same way.

The golden rule of tables — the two bars and the roles of the columnsThe order of the columns to the left of the gold bar hierarchizes the data, the columns between the gold and blue bars are display columns, and the columns to the right of the blue bar are aggregate values. Dragging columns to reorder them is itself the analysis operationLeft of the gold bar: grouping columns (order = hierarchy)Gold barBetween the bars: display columnsBlue barRight of the blue bar: aggregate values (Sum, %, and so on)Dragging a column = an analysis operation (Process → Stack gives a per-process stack aggregation)

4.2. Narrow to the Interval When the Issue Occurred

Drag on the graph to select a range, then right-click and choose “Zoom,” and the aggregation switches to that interval only. The principle of performance investigation is to always look only at “the interval when the issue was occurring” (Chapter 9).

4.3. Load Symbols and View Stacks by Function Name

To read stacks by function name, run Trace > Load Symbols from the menu.14 By default it refers to Microsoft’s public symbol server (msdl.microsoft.com), so stacks in Windows itself resolve as long as you have an Internet connection.

To see function names in your own app as well, add the folder containing your app’s PDBs in Trace > Configure Symbol Paths.8 What a PDB is, and why you should always keep one even for release builds, is summarized in “What Is a PDB (Program Database)?.”

For .NET, Treat NGen Images and JIT Code Separately

For .NET Framework NGen native images, WPR generates NGen PDBs (.ngenpdb) at capture time, places them in a folder next to the trace, and WPA refers to them automatically.8 This mechanism is specific to NGen images; your own code in an ordinary .NET app that runs under the JIT is not covered.

The mapping from JIT code addresses to function names is resolved from the JIT events the CLR emits. So when you investigate a .NET app, prepare a recording profile (.wprp) that enables the CLR providers (Microsoft-Windows-DotNETRuntime and its Rundown counterpart), combine it in the form wpr -start GeneralProfile -start MyDotNet.wprp!ProfileName, just as with your own provider in Chapter 2, so that the CLR events are included in the trace (you can check which built-in profiles your local WPR offers with wpr -profiles).

On top of that, keep the PDBs generated by the build for mapping to source lines, and add them to the symbol path described above.

Symbol resolution for reading stacks by function nameWhen you run Trace Load Symbols, Windows itself resolves from Microsoft's public symbol server and your own app from the build PDBs added to the symbol path. NGen images resolve from the NGen PDBs WPR generates, and JIT .NET code from the CLR JIT events in the trace plus the build PDBsTrace > Load SymbolsWindows itself: public symbol server (msdl)Your own app: build PDBs added in Configure Symbol Paths.NET Framework NGen images: .ngenpdb generated by WPRJIT .NET code: CLR JIT events in the trace + build PDBs

4.4. Decide Whether to Investigate CPU, Waits, or I/O

Once you are set up, look at whether CPU was high or low during the interval of the issue. If it was high, go to Sampled in Chapter 5. If it was low, first check whether a particular core or thread is pinned; if not, follow the waits with Precise in Chapter 6. If the disk is suspected, go to Chapter 7.

Choosing a WPA graph from the symptomZoom to the interval of the issue; if CPU is high go to CPU Usage Sampled; if it is low but slow, check for a pinned core and then do wait analysis in CPU Usage Precise; if the disk is suspected go to Disk Usage and File IOHighLow but slowYesNoDisk suspectedZoom to the interval of the issueCPU in that interval?Chapter 5: CPU Usage (Sampled) for 'who is burning CPU in which function'One core or one thread pinned?Chapter 6: CPU Usage (Precise) for 'what it was waiting for'Chapter 7: Disk Usage / File I/O to identify the culprit

5. When CPU Is High — “Who Is Burning CPU in Which Function” With CPU Usage (Sampled)

What this chapter looks for is the call path that uses a large share of CPU time.

If the CPU is pinned, what you look at is CPU Usage (Sampled). This is sampling data that records, about every millisecond on every CPU, “which process’s stack is running right now,” and the ratio of sample counts is directly the breakdown of CPU time.6

How CPU Usage Sampled worksAbout every millisecond, the stack running on every CPU is recorded, and the ratio of aggregated samples is the breakdown of CPU time. Read down from process to thread, stack, and function. Short activity that finishes between samples is not capturedInterrupt about every 1 msRecord the 'stack running right now' on every CPUAggregate the samples (ratio = breakdown of CPU time)Drill down Process → Thread → Stack → functionShort activity that finishes between samples is not captured

Drill Down From Process to Stack to Function

  1. From Computation in Graph Explorer, place CPU Usage (Sampled) on the Analysis tab and select the Utilization by Process, Stack preset.5
  2. Look at the processes in descending order of Weight (or Count). What was “50%” in Task Manager is first identified at the process level.
  3. Expand the Stack column of the culprit process. Stacks are aggregated as a tree, and descending along the path where the numbers do not drop much at each branch brings you to the function that is burning CPU. If symbols have resolved, it is a straight line to the exact function in your own code.
  4. If expanding the tree is tedious, switch the graph display to Flame (flame graph). Width is drawn as the share of CPU time, so which call path dominates is obvious at a glance. CPU Usage (Sampled) also comes with a Flame by Process, Stack preset.13

Sampled Cannot Measure the Exact Duration of a Single Run

Because it is sampling, short activity that finishes between samples is not captured.6 Remember that it is a tool for seeing “where CPU was used in total,” not a tool for measuring the exact execution time of each individual run.

6. When CPU Is Low but It Is Still Slow — CPU Usage (Precise) and Wait Analysis

6.1. Before Wait Analysis, Check for a Single Pinned Core

“Overall CPU usage is low” does not mean “the CPU is not the bottleneck.” On a 16-core PC, serial work pinned to one core (a single UI thread running flat out) shows up as only about 6% overall.

First check, with Sampled from Chapter 5 or with Utilization by CPU in CPU Usage (Precise), whether a particular core or thread is pinned. If nothing is pinned either, assume the work is not unable to use the CPU but is waiting for something, and proceed to wait analysis. From here on, the tool is CPU Usage (Precise).

6.2. Separate “Time Spent Waiting” From “Waiting for a CPU After Waking”

Where Sampled is sampling, Precise is a complete record of context switches (thread switches). A thread enters a wait, is woken by someone (Ready), and gets onto a CPU — each such round trip is kept as one row, and you can read the following columns.74

Column Meaning
NewThreadStack On which stack the thread entered the wait (= what it was doing when it stopped)
Waits (us) Time spent waiting
Ready (us) Time it had to wait between being woken and getting onto a CPU (CPU contention)
ReadyingProcess / ReadyingThreadId The process and thread that woke it (released the wait)
ReadyThreadStack On which stack the waking side woke it
One wait round trip and the corresponding columnsA thread enters a wait on the stack kept in NewThreadStack and waits for the Waits time. When someone wakes it, that party is kept in ReadyingProcess and ReadyThreadStack, and the thread waits the Ready time for CPU contention before running againEnters a wait (kept in NewThreadStack)Someone wakes it (ReadyingProcess / ReadyThreadStack)Gets onto a CPURunningWaiting (Waits (us))Ready (waiting for CPU contention)Running again

6.3. From the Delayed Thread, Trace Each Waker in Turn

The reading order is as follows.4

1. Set Up the Graph and Columns

Apply the Utilization by Process, Thread preset and add NewThreadStack and ReadyThreadStack to the columns.

2. Select the Thread of the Delayed Operation

First identify the thread that was executing the delayed operation, such as the UI thread or the thread handling the request in question. If you only look in descending order of total Waits, healthy threads that wait on purpose, such as a message pump or a timer, come out on top.

If the target thread’s CPU Usage (ms) is large, it is a CPU problem from Chapter 5; if Waits dominate, it is a wait problem.

3. See “What It Was Doing When It Stopped” in NewThreadStack

Expand NewThreadStack. WaitForSingleObject or EnterCriticalSection means a lock wait; inside synchronous I/O such as ReadFile, an I/O wait; inside a socket receive, waiting for the peer to respond.

4. See “Who Released the Wait” in ReadyThreadStack

Expand ReadyThreadStack and check ReadyingProcess / ReadyingThreadId. If it was woken from the kernel’s KiTimerExpiration, it was a timer (it slept until the timeout); if it was woken from I/O completion processing, that confirms it was an I/O wait.4

5. Repeat the Same Investigation on the Waker

If the waker is another thread or another process, investigate that thread next with the same procedure.

For example, a chain such as A was waiting for B to release a lock, B was waiting for an RPC response from C, and C was waiting for disk I/O. Following this chain to its root gives you the critical path of the delay.7

The critical-path chain traced in wait analysisSee in delayed thread A's NewThreadStack what it was doing when it stopped, identify the waker B from ReadyThreadStack and ReadyingProcess, investigate B with the same procedure, and trace down to the root disk I/ONewThreadStack: stopped in a lock waitNewThreadStack: waiting for an RPC responseNewThreadStack: waiting on synchronous I/OCompletion wakes CThe response wakes BReleasing the lock wakes A (appears in ReadyThreadStack)Thread A (the delayed operation)Thread B (holding the lock)Process CDisk I/O (the root)

Turn the Cause of the Wait Into a Design Review

In a project where “we made it multithreaded but it did not get faster,” for example, this procedure shows every worker lined up behind a single lock exactly as it is. Avoiding lock contention by design is discussed in “Practical Multithreading Best Practices: .NET Edition,” and the Windows mechanism that runs on completion notifications instead of waiting in synchronous I/O in “I/O Completion Ports (IOCP) and the .NET Thread Pool.” Pinning down “who it was waiting for” in WPA and then fixing it with these design principles is one continuous flow.

7. Disk and File I/O — Identifying “Someone Is Scanning the Disk”

The classic culprit behind “the whole PC is slow” is not the CPU but the disk. Investigate with Disk Usage and File I/O in the Storage category.15

7.1. Separate “Device Processing” From “Queueing” in Disk Usage

Disk Usage is the record of disk I/O. Start by distinguishing the following two columns.6

Column Time it represents
Disk Service Time The time the disk device actually took to process that I/O
IO Time The time from when the I/O entered the OS queue until it completed

IO Time is always at least Service Time, the difference being the time spent in the queue. If IO Time is much longer than Service Time, that I/O was waiting in the queue.6

Do not, however, conclude the cause from the time difference alone. Whether another process built the queue, or the process’s own heavy I/O lined up on a slow device, is not yet decided. Look at the device’s own response in Service Time, then settle it with the breakdown by process, path, and stack.

7.2. Which Process Issued I/O to Which File

So, with the Utilization by Process, Path Name, Stack preset, look at which process issued I/O to which file from which stack, in descending order of IO Time or Size.15 The two answers that come up most often in the field are these.

  • Antivirus software was scanning every file. It shows up as the antivirus process issuing a large volume of reads during the interval when the app was slow to start. The process name, path, and volume are evidence ready to use in a discussion about exclusion settings.
  • Another process was writing heavily. Backup, an indexer, excessive logging, and so on. When a write reaches the disk involves the cache manager, so the fact that “the moment of the write” and “the moment the disk is busy” can diverge is also explained in “Cache Manager — When Does Your WriteFile Actually Reach the Disk?.”

7.3. See Slowness Before the Disk With File I/O

File I/O is one layer up: the record of file operations the app issued (Create, Read, Write, and so on), and presets such as Duration by Process, Thread, Type aggregate time per file name and per operation.15

A case where time is spent in the file system or a filter driver before reaching the disk does not appear in Disk Usage, so the discrepancy itself, “Disk Usage is quiet but File I/O is slow,” is a clue. If you want to start from how synchronous and asynchronous I/O work, see “Synchronous and Asynchronous I/O — What OVERLAPPED Really Means.”

The different layers seen by File IO and Disk UsageAn app's file operation passes through the file system and filter drivers and reaches the disk device from the OS IO queue. File IO records the operations at the upper layer, Disk Usage records the IO that reached the disk, and the difference between IO Time and Disk Service Time shows the time in the queueApp: ReadFile / WriteFileFile system and filter drivers (the layer File I/O sees)OS I/O queueDisk device (the layer Disk Usage sees)Time spent here does not appear in Disk UsageIO Time − Disk Service Time = time in the queueDisk Service Time = device processing time

7.4. When You Suspect a Memory Shortage, Check Physical Memory Too

The line of reasoning “maybe it is short of memory and swapping” can be given a first triage in Task Manager and Resource Monitor before you move on to WPA.

Do not rule out a memory shortage by looking at committed memory alone. Even with headroom in commit, physical memory pressure can trim working sets and keep hard faults coming. Also check available physical memory and “Hard Faults/sec” in Resource Monitor. For how to read them, see the article What Does Windows’ “Memory Usage” Actually Mean?.

8. Slow Boot and Logon — The Entrance to Boot Tracing

With the “boot takes 3 minutes” type, the problem is over before you can run wpr -start by hand. WPR has a boot trace, which you can set up so that the OS starts recording automatically on the next boot.3

Set Up Recording for the Next Boot, Then Save After the Restart

As in Chapter 3, decide the capture approval and ETL handling, then operate from a Command Prompt opened as administrator.

:: 1. Set up automatic recording on the next boot
wpr -boottrace -addboot GeneralProfile -filemode

:: 2. Restart (reproduce the slow boot)

:: 3. After boot, stop the recording and save (this also clears the setup)
mkdir C:\temp 2>nul
wpr -boottrace -stopboot C:\temp\boot.etl "Boot takes 3 minutes"
Boot trace flowSet up automatic recording on the next boot with addboot, then restart; the OS starts recording automatically at boot. Saving with stopboot after logon also clears the setup. To abandon it, clear it with cancelbootSet up with wpr -boottrace -addbootRestart (reproduce the slow boot)The OS starts recording automatically at bootSave with -stopboot after logon (also clears the setup)To abandon, -cancelboot

After Capturing, the Same “CPU, Wait, or I/O” Triage

The boot and shutdown measurement that xbootmgr once handled can also be run in current WPR with options such as -onoffscenario Boot.3

The captured trace is read with the same toolkit as the previous chapters. Look at which process was created when, in chronological order, in the Processes graph; zoom to the interval where boot is stuck; and classify it as CPU, wait, or disk. Patterns such as startup apps waiting for something in series, or a service start stalled on a particular I/O, come into view.

Boot analysis is a deep specialty in its own right, so this article goes only as far as the entrance: “an issue you cannot catch by hand can still be captured with WPR.” Start by getting the overall picture with a GeneralProfile boot trace.

9. The Working Pattern — Classify, Zoom, Stack, Repeat

Now that the tools are clear, here is the pattern for the investigation as a whole. The order pin down the time and narrow the interval, classify, then trace the stack is the same for every symptom.

From Pinning Down the Time to Verifying the Hypothesis

  1. Pin down the time of the phenomenon. Not “it was slow” but “it was slow from 10:23:40 to 10:24:10.” App logs, the event log, a note from the person who was operating it — anything will do. If your own app writes milestones to ETW or the event log, the events inside the trace serve directly as time markers.
  2. Zoom to that interval only. An aggregation over the whole trace is smoothed into averages, and the anomaly that matters gets diluted. WPA analysis is always a comparison of “the interval that was abnormal” against “the interval that was normal.”
  3. Classify “CPU, wait, or I/O” first. Look at CPU Usage (Sampled); if it is burning, go to Chapter 5. If it is not burning, go to Waits in CPU Usage (Precise) (Chapter 6). If IO Time in Disk Usage is inflated, go to Chapter 7. Passing through this three-way fork first keeps you from getting lost.
  4. Repeat hypothesis → zoom → stack. If you think “antivirus?”, narrow to that process and confirm it with the stack. If it does not hold, move to the next hypothesis. Not drawing a conclusion before you have gone down to the stack and confirmed it is the discipline of this kind of investigation.
The iterative loop of a performance investigationPin down the time of the phenomenon, zoom to the interval, classify CPU, wait, or IO, form a hypothesis and narrow down, and confirm with the stack. If it holds the cause is confirmed; if not, repeat with the next hypothesisConfirmedDid not holdPin down the time of the phenomenonZoom to that intervalClassify CPU, wait, or I/OForm a hypothesis and narrow downConfirm with the stackCause identified → fix it

Decide How to Handle the ETL Before Capturing

An ETL file captures a wide view of the inside of the system: the names of every process, the paths of files that were opened, the modules that were loaded, and (depending on the profile) registry key names.

A standard GeneralProfile capture does not include data bodies such as communication contents, but if you enabled a custom provider, the payload of its events (such as strings the app recorded) goes in as is.

Confirm what the providers you enabled emit, and then treat it as a file confidential enough to warrant care before it leaves the company. As with packet capture, build minimal capture, agreement with the recipient, and a retention period and deletion into the procedure.

What an ETL file captures and how to handle itAn ETL captures every process name, the paths of opened files, modules, and depending on the profile registry key names; enabling a custom provider adds its payload too. Treat it as confidential, and build minimal capture, agreement with the recipient, and a retention period and deletion into the procedureETL fileEvery process name, file paths, modules(Depending on the profile) registry key namesCustom provider payloads (strings the app recorded)Treat as confidential: minimal capture, agreement with the recipient, retention period and deletion

10. Summary

A WPR/WPA investigation can be organized into the following three stages.

  1. Capture the interval you need. Capture with wpr.exe, which ships with Windows 8.1 and later, and read the ETL in WPA on your own machine. The basics are wpr -start GeneralProfile -filemode → reproduce → wpr -stop trace.etl. If you can reproduce it, File mode within a few minutes; if you wait for it, Memory mode. For slowness during boot, use wpr -boottrace. Longer is not better, and the internal information in the ETL is treated as confidential.
  2. Narrow the time range and choose the direction of investigation. In WPA, left of the gold bar is grouping and right of the blue bar is aggregate values. Master column order, zooming to the interval, and symbol configuration, then isolate CPU, waits, or I/O. Provide PDBs for your own app, and for .NET JIT code check the CLR events at capture time as well.
  3. Trace down to the stack and verify the hypothesis. If CPU is high, drill down from process to function in Sampled. Even if overall CPU is low, first check for a single pinned core or thread, and if there is none, follow the waits in Precise. For disk, do not decide the cause from the time difference alone; look at the device’s response and the breakdown of who issued the I/O.

What you trace in wait analysis is NewThreadStack (what it was doing when it stopped) → Waits (how long it waited) → ReadyingProcess and ReadyThreadStack (who woke it). Investigate the waker the same way, and you can follow the chain of delays to its root.

WPA may overwhelm you at first with the amount of information on screen. Still, once you hold on to “column order is how data is aggregated” and “Sampled is where CPU was used, Precise is who it waited for,” the rest is the repetition of pinning down the time, zooming, classifying, and checking the stack. The next time a consultation comes in saying “CPU has headroom but it is still slow,” capture a trace in this order and read through it.

KomuraSoft LLC handles investigations of system-wide performance problems such as “the whole PC became slow and we cannot find the cause,” “CPU has headroom but the app is slow,” and “only one particular environment is extremely slow to boot.” We handle everything as one continuous engagement, from capture design with WPR/WPA (which environment, which profile, and how much to capture) through trace analysis to fixing the app-side or settings-side cause.

References

  1. Microsoft Learn, Introduction to WPR. That WPR is an ETW-based performance recording tool; that the command-line edition WPR.exe ships with Windows 8.1 and later with no additional install; its relationship to the GUI edition WPRUI.exe; and the concept of recording profiles. ↩ ↩2 ↩3

  2. Microsoft Learn, Windows Performance Analyzer. That WPA is included in the Windows ADK, is an analysis tool that builds graphs and data tables from ETW events recorded by WPR, Xperf, and the like, and can open and analyze any ETL file. ↩ ↩2 ↩3 ↩4

  3. Microsoft Learn, WPR Command-Line Options. The syntax of wpr -start/-stop/-cancel/-status/-profiles; -filemode (the default is memory mode); specifying several profiles at once; boot traces with -boottrace (addboot/stopboot/cancelboot); and recording On/Off transitions such as Boot with -onoffscenario. ↩ ↩2 ↩3 ↩4 ↩5 ↩6

  4. Microsoft Learn, CPU Analysis. Definitions of the CPU Usage (Precise) graph columns (NewThreadStack, ReadyThreadStack, ReadyingProcess, Waits, and the like); the procedure of expanding ReadyThreadStack and following ReadyingProcess/ReadyingThread to the root cause of a wait; and how to tell a wake from KiTimerExpiration (a timer wait) from one caused by I/O completion. ↩ ↩2 ↩3 ↩4 ↩5

  5. Microsoft Learn, Troubleshoot processes and threads by using WPR and WPA. Configurations such as reading CPU Usage (Sampled) as Process→Stack under high CPU usage, and using CPU Usage (Precise) Readying Process, Readying Thread, Readying Stack, and the Wait column in wait analysis; and the correspondence table of profiles and graphs by symptom. ↩ ↩2

  6. Microsoft Learn, Exercise 2 - Evaluate Fast Startup Using Windows Performance Toolkit. That CPU Usage (Sampled) is sampling at about a 1-millisecond interval and short activity between samples is not recorded; the procedure of drilling down process → thread → stack to identify the breakdown of CPU consumption; and the meaning of Disk Usage IO Time (including queue time) and Disk Service Time (disk processing time). ↩ ↩2 ↩3 ↩4 ↩5

  7. Microsoft Learn, Exercise 3 - Understand Critical Path and Wait Analysis. The concept of critical-path analysis (the Running, Ready, and Waiting classification); the meaning of the CPU Usage (Precise) table columns NewThreadStack, ReadyThreadStack, ReadyingProcess, Waits, Ready, and the like; and the procedure of following each waking thread in turn to unravel a chain of delays. ↩ ↩2 ↩3

  8. Microsoft Learn, Loading Symbols. That when _NT_SYMBOL_PATH is unset WPA refers to Microsoft’s public symbol server (msdl.microsoft.com) by default; adding PDB paths for your own components; and that WPR generates PDBs for .NET managed symbols in a .ngenpdb folder next to the trace and WPA refers to them automatically. ↩ ↩2 ↩3

  9. Microsoft Learn, Built-in Recording Profiles. The list of recording profiles built into WPR (CPU usage, Disk I/O activity, File I/O activity, Registry I/O activity, Networking I/O activity, and others) and what each profile records. ↩

  10. Microsoft Learn, Logging Mode. That the recording modes are File (a continuous file) and Memory (an in-memory circular buffer) and the default is Memory; that Memory suits an issue whose timing is unknown and older events are overwritten; and that File’s only ceiling is free disk space and a file that is too large may not be analyzable in WPA. ↩ ↩2

  11. Microsoft Learn, WPR How-to Topics. The procedure for starting and stopping a recording in WPRUI; choosing a profile, detail level, and Logging mode; and the caution that a long recording can make the file huge and unanalyzable in WPA, so Memory mode should be chosen. ↩ ↩2

  12. Microsoft Learn, Graph Explorer. That the Graph Explorer window lists graph thumbnails in categories such as System Activity, Computation, Storage, and Memory; and that you drag a graph onto the Analysis tab to display it together with a table. ↩

  13. Microsoft Learn, Graphs (WPA Features). WPA’s Flame graph display; the table structure in which columns to the left of the gold bar are grouping and columns to the right of the blue bar are aggregate values; and the CPU Usage (Sampled) Flame by Process, Stack preset. ↩ ↩2

  14. Microsoft Learn, Load Symbols or Configure Symbol Paths. Loading symbols with Load Symbols from WPA’s Trace menu; and the procedure for setting and changing symbol paths in the Configure Symbol Paths dialog. ↩

  15. Microsoft Learn, List of WPA Graphs. The list of graphs available in WPA. Disk Usage presets such as IO Time by Process, IO Type; Service Time by Process, Path Name, Stack; Utilization by Process, Path Name, Stack; and File I/O presets such as Duration by Process, Thread, Type. ↩ ↩2 ↩3

Recent articles sharing the same tags. Deepen your understanding with closely related topics.

These topic pages place the article in a broader service and decision context.

This article connects naturally to the following service pages.

Frequently Asked Questions

Common questions about the topic of this article.

Where can I get WPR and WPA? Can they be used in a customer environment where nothing can be installed?
The capture tool wpr.exe (the command-line edition) ships with Windows 8.1 and later, so it can be used without any additional install. The GUI edition, WPRUI, and the analysis tool WPA (Windows Performance Analyzer) are included in the Windows ADK (Windows Assessment and Deployment Kit) and require a separate install. In practice, if you split the work as "capture an ETL file with only the OS-standard wpr.exe in the customer environment, take it back, and analyze it in WPA on your own machine," you can investigate system-wide performance even at a site where no software can be added.
Why is it slow even though Task Manager shows CPU headroom? What can WPA tell me?
When it is slow even though CPU usage is low, the work is not unable to use the CPU; it is stopped "waiting for something." Typical cases are contention for a lock, waiting for synchronous I/O to complete, and waiting for another process to respond. Task Manager shows only the result, the usage figure, but CPU Usage (Precise) in WPA shows, from a per-context-switch record, where a thread started waiting (NewThreadStack), how long it waited (Waits), and who woke it (ReadyingProcess and ReadyThreadStack). By tracing each party that made it wait in turn, you can identify "the culprit of the slowness" at the function level.
How long should I capture a trace? Won't the file become huge?
If you can reproduce the issue, the basic approach is to start just before the reproduction, stop just after, and keep it within a few minutes. WPR's default is Memory mode, which records to an in-memory circular buffer; older events are overwritten first, so it suits waiting for an issue whose timing is unknown. File mode, enabled with -filemode, keeps everything in a continuous file, but its only ceiling is free disk space, and a file that is too large may not be analyzable in WPA. Use Memory mode for a long wait and File mode for a short, reliable reproduction.
When should I use PerfView, and when WPA?
Both tools handle ETW traces, but their strengths differ. PerfView has a deep understanding of the .NET runtime and is strong at investigations specific to managed apps, such as GC, allocations, and JIT. WPA suits reading OS-wide CPU, disk, file I/O, power, and more across graphs and tables, and is the first choice when "it is not a specific app but the whole PC that is slow," "multiple processes are involved," or "something outside the app (antivirus, a driver, another process) is suspected." As a rule of thumb: slowness in your own .NET app alone, PerfView; slowness across the whole system, WPR/WPA.
Is it safe to run WPR in a customer's production environment?
A short capture is common in practice, but it is not unconditionally safe. ETW is lightweight, but it records a large volume of events with stacks, so it consumes a certain amount of CPU and memory. Put considerations such as starting just before the reproduction step and stopping just after, keeping the capture within a few minutes, and running it at a time with little impact on business through the same approval process as any ordinary change. Also, an ETL file contains internal system information such as process names, file paths, and executable details, so you should also decide in advance how it will be handled if it is taken outside the company (minimization, retention period, deletion).

Author Profile

Profile page for the article author.

Go Komura

Representative of KomuraSoft LLC

Focused on Windows software development, technical consulting, and investigations into failures that are difficult to reproduce.

Back to the Blog