The Win32 Thread Pool API — Concurrency That Does Not Create Threads, with CreateThreadpoolWork
· Updated: · Go Komura · Windows, Multithreading, C++, Windows Development, Win32 API, Performance Improvement
Revision history (1 updates, last updated Sep 7, 2026)
A log of the changes made to this article. Where a pre-update version was archived, it stays readable at a permanent DOI link.
- Retranslated as a full translation of the current Japanese original. The previous English version was an abridgement that dropped subsections, tables, diagrams, and paragraphs; all of them have been restored to match the Japanese article, and the knowledge map section has been added where the Japanese article has one. The technical claims are the same as in the Japanese version. Read the version before this update (DOI: 10.5281/zenodo.22170915)
- First published
Cite this article(DOI: 10.5281/zenodo.22170914)
This article is archived on Zenodo. Below are both the DOI that always resolves to the latest version and the DOI pinned to the version you are reading.
Go Komura (2026). The Win32 Thread Pool API — Concurrency That Does Not Create Threads, with CreateThreadpoolWork. KomuraSoft LLC. https://doi.org/10.5281/zenodo.22170914 https://comcomponent.com/en/blog/win32-thread-pool-api/
- DOI (latest version)
- 10.5281/zenodo.22170914
- DOI (this version)
- 10.5281/zenodo.22615677
“One CreateThread per client.” “One thread for the timer.” “One for waiting on an event.” In native Windows code, the reflex is to add a thread for every small job. Each thread consumes a stack and a kernel object, and creating and destroying it costs something as well.
What absorbs this mismatch is the OS-standard Win32 thread pool API. The app hands over the processing it wants to run as a callback and leaves worker-thread management to the OS. “Not creating threads” does not mean that threads become unnecessary; it means that the app does not create and destroy one itself for every job.1
This article is for developers who write Windows apps, services, and DLLs in C/C++, and proceeds in the order choosing the tool → selecting the object → implementing work → cautions for callbacks and shutdown. The subject is the CreateThreadpoolWork family of APIs redesigned in Windows Vista. The code is an excerpt that shows the skeleton of usage; app-specific parts such as the queue are omitted.
1. The Bottom Line First
If you process a large number of short-lived jobs, manage jobs rather than threads. The app, however, still designs when a job stops and when it can be released.
There are three axes of judgment.
- Choose the tool. If standard C++ or .NET tools are enough, use them. This API’s turn comes when you want to unify timers, waits, and I/O completion in Win32, or when you want to split pools.
- Hand over work in units of jobs. Use the new API’s work, timer, wait, and io, and leave jobs that need thread-specific state, such as a thread priority or a COM STA, on dedicated threads.
- Build shutdown as part of the set. The basic form is “stop submitting → wait for completion → close”. Inside a callback, do not block for a long time, do not synchronously wait for work on the same pool, and restore the thread’s state.
If you want to get something running first, start with the work example in Chapter 4 and then check the callback discipline in Chapter 5. Bulk shutdown of multiple objects is covered in Chapter 6, and DLL unloading in Chapter 7.
2. Choosing Between a Pool, a Dedicated Thread, and the Standard Library
2.1 A Pool Suits Short-Lived, High-Volume, Wait-Centric Work
A thread pool is a set of worker threads managed by the OS. The workers execute callbacks one after another, and the OS adjusts their number according to load. By handing over the work you used to do by creating and tearing down your own threads, you reduce both management code and the cost of creation and destruction.1
The official documentation lists as candidates apps that issue a large number of small work items in parallel, apps that frequently create and destroy short-lived threads, apps that process independent work in parallel in the background, and apps that have dedicated threads waiting on kernel objects or events. Search, network I/O, and consolidating waiter-only threads are typical.1
On the other hand, work that needs a change of thread priority, needs a COM STA, or keeps running for the whole lifetime of the process stays on a dedicated thread. The criterion is whether it is enough for the processing to run, or whether the thread itself needs a “personality”.
flowchart TB
accTitle: Choosing between a dedicated thread and a pool
accDescr: First check whether the thread needs a personality such as a priority or an STA and whether it runs for a long time; only short-lived, high-volume, or wait-style work that matches neither goes on the thread pool
q1{"Needs a personality such as priority or STA?"} -->|"Yes"| ded["Keep on a dedicated thread"]
q1 -->|"No"| q2{"Runs for a long time?"}
q2 -->|"Yes"| ded
q2 -->|"No"| pool["Put on the thread pool"]
pool -.-> ex["Short-lived work, waits, timers, I/O completions"]
Figure 1: Only work that “needs no personality and finishes quickly” belongs on the pool. Everything else stays on a dedicated thread as before.
2.2 First Ask Whether Standard C++ or .NET Is Enough
If the granularity is covered by C++ std::async or std::thread, the standard library is the first candidate. It is portable, and the behavior of std::async and future can be handled on the basis of the standard.2
The reasons to use the Win32 thread pool directly are requirements such as wanting a unified callback mechanism that includes timer, wait, and io; wanting separate pools or thread counts per kind of work; or not wanting to own threads inside a DLL or COM component. Separate the question of whether you simply want parallelism from whether you need Windows-specific control.
flowchart TB
accTitle: Deciding which layer's tools to write with
accDescr: If standard C++ async or thread is enough, use it; use the Win32 thread pool directly when you need to unify timers, waits, and I/O completion, split pools or control thread counts, or avoid owning threads inside a DLL or COM component
q1{"Are the standard C++ tools enough?"} -->|"Yes"| std["std::async / std::thread"]
q1 -->|"No"| q2{"What do you need?"}
q2 -->|"Unify timer, wait, and io"| tp["Win32 thread pool"]
q2 -->|"Split pools / control counts"| tp
q2 -->|"Avoid own threads in a DLL"| tp
Figure 2: When in doubt, start with the standard library; this API’s turn comes when a requirement it cannot express appears.
In .NET, ThreadPool and Task play the same role, and I/O completion is tied to IOCP. For a detailed explanation of the relationship, see the article on IOCP and the .NET thread pool. There is no need to go down to the Win32 API where the tools in the layer above are enough.
2.3 In New Code, Use the New API from Vista Onward
There are two generations of Win32 thread pool API: the legacy API that has continued since Windows 2000, such as QueueUserWorkItem and RegisterWaitForSingleObject, and the new CreateThreadpoolWork family that was completely redesigned in Windows Vista.
The new API unified the kinds of worker thread and provided a single timer queue, dedicated persistent threads, multiple independent pools within a process, and cleanup groups. The official documentation also cites the new API’s simplicity, reliability, performance, and flexibility as its advantages.13
The legacy API also has a structural constraint: there is no way to cancel work once it has been queued. Use the new API in new code, and when reviewing existing code, start from the following correspondence.3
flowchart TB
accTitle: Correspondence between the legacy thread pool API and the new API
accDescr: The legacy API's QueueUserWorkItem is replaced by the new API's work object, timer queues by timer, registered waits by wait, and BindIoCompletionCallback by io
o1["QueueUserWorkItem"] --> n1["work"]
o2["Timer queues"] --> n2["timer"]
o5["Registered waits"] --> n3["wait"]
o4["BindIoCompletionCallback"] --> n4["io"]
Figure 3: The migration target from the legacy API is fixed one-to-one. An inventory of existing code can start from this correspondence table.
3. How It Works — Four Objects with Different Firing Conditions
3.1 Choose by What Should Trigger the Callback
At the center of the new API are the following four kinds. Choose not only by “what to process” but by what should trigger the callback.4
| Object | Creation function | Callback firing condition |
|---|---|---|
| work | CreateThreadpoolWork |
When submitted with SubmitThreadpoolWork |
| timer | CreateThreadpoolTimer |
When the specified time or period arrives |
| wait | CreateThreadpoolWait |
When a kernel object becomes signaled |
| io | CreateThreadpoolIo |
When asynchronous I/O on the associated handle completes |
The firing conditions differ, but the workers of the same pool do the executing. Instead of writing periodic processing, event response, and I/O completion handling each on its own dedicated thread, you can align them on a common callback mechanism.
flowchart TB
accTitle: Separating firing conditions from callback execution
accDescr: The pool's workers execute callbacks that objects with firing conditions have made ready, and a worker that has finished is used for the next callback as well, so the app does not create a thread per job
app["App specifies the job and its firing condition"] --> obj["Object waits for the condition"]
obj --> ready["Callback becomes ready to run"]
ready --> workers["Pool workers execute it"]
workers --> done["Finish and return the worker"]
done --> next["Reused for the next callback too"]
Figure 4: Separating the mechanism that waits for a trigger from the workers that execute the processing reduces per-job thread management.
3.2 “Threads That Only Sleep and Wait” Can Be Replaced Too
The pool’s benefit is not limited to CPU-bound work. Timers are consolidated into a single timer queue, and waits on multiple handles onto a small number of wait threads.1
For example, if you have five threads that sleep only so they can “run when the event is signaled”, you can think of replacing them with five wait objects. Instead of holding a thread that sleeps on its own for each one, you run the on-signal processing as a callback.
flowchart TB
accTitle: Replacing waiter-only threads with wait objects
accDescr: Waiter-only threads that slept one per event are consolidated onto the pool's wait threads when replaced with wait objects, and the callback runs only when signaled
old2["5 waiter-only threads sleeping individually"] -.-> waste["Consumes 5 stacks and 5 threads"]
new2["5 wait objects"] --> agg["Consolidated onto the pool's wait threads"]
agg --> cb2["Callback runs only on signal"]
Figure 5: “Threads that only sleep and wait” can be eliminated by turning them into wait objects. This is the obvious first move in a pool migration.
4. Implementation Basics — Create a work, Submit It, and Shut Down Safely
4.1 First, One Round Trip from Creation to Shutdown
Walk through usage once with the most basic object, work. CreateThreadpoolWork binds the callback to a context, and SubmitThreadpoolWork requests execution. If the third argument is NULL, the process’s default pool is used. Many uses are fine with this default pool.56
The following is an excerpt that shows the procedure, not a complete program that compiles as is. WORK_QUEUE, ITEM, Enqueue, Dequeue, and ProcessItem stand for app-side processing. Implement the queue’s mutual exclusion, stopping the submitter, and failure handling separately, and if creating the work fails, do not proceed to the subsequent submit, wait, and release.
VOID CALLBACK WorkCallback(PTP_CALLBACK_INSTANCE instance,
PVOID context, PTP_WORK work)
{
// The context is fixed at creation time. Per-item data is passed through a synchronized queue
WORK_QUEUE* queue = (WORK_QUEUE*)context;
ITEM* item = Dequeue(queue); // Take one item under mutual exclusion
ProcessItem(item);
}
// 1) Create (bind the callback to the shared context, i.e. the queue)
PTP_WORK work = CreateThreadpoolWork(WorkCallback, &queue, NULL);
if (!work) { /* Failure handling with GetLastError */ }
// 2) Submit once for each item you enqueue (keep the item count and the submit count equal)
Enqueue(&queue, item);
SubmitThreadpoolWork(work);
// 3) Stop the submitter, then wait for completion (TRUE also attempts to cancel work that has not started)
WaitForThreadpoolWorkCallbacks(work, FALSE);
// 4) Close
CloseThreadpoolWork(work);
flowchart TB
accTitle: Lifecycle of a work object
accDescr: Create with CreateThreadpoolWork; submitting with SubmitThreadpoolWork runs callbacks in parallel. At shutdown, first stop new submits, wait for all callbacks to complete with WaitForThreadpoolWorkCallbacks, then close with CloseThreadpoolWork
c["Create with CreateThreadpoolWork"] --> s["Submit with SubmitThreadpoolWork (repeatable)"]
s --> run["Callbacks run in parallel"]
run --> stop3["Stop new submits"]
stop3 --> w["Wait for completion with WaitForThreadpoolWorkCallbacks"]
w --> cl["Close with CloseThreadpoolWork"]
Figure 6: The shutdown order is “stop submitting → wait for completion → close”. Skip any step and you get use-after-free or a race.
4.2 A work Can Be Reused, but the context Does Not Change per Submit
The same work object can be submitted multiple times, even before the previous callback has finished. Each submit runs the callback, and multiple instances run in parallel. The number of threads actually used may be adjusted by the pool for efficiency.7
What to distinguish here is the work from the data of each work item. The context passed to the callback is fixed at creation time. To process N items of different data with one work, make a synchronized queue the context, as in the example above. Submit once for each item you enqueue, and have the callback take one item from the queue. A design that creates a work object per item is also fine.5
flowchart TB
accTitle: Receiving per-item data through a fixed context
accDescr: The context is fixed to a shared queue when the work is created; the submitter submits once for each item it enqueues, and each callback running in parallel takes one item at a time from the queue under mutual exclusion
create["Fix the context at work creation"] -.-> queue["Shared queue with mutual exclusion"]
item["Enqueue one item"] --> queue
item --> submit["Submit once per item"]
submit --> callbacks["Callbacks run in parallel"]
callbacks --> dequeue["Dequeue one item each"]
queue --> dequeue
dequeue --> process["Process the item taken"]
Figure 7: What is fixed is the context that points to the queue; per-item data is passed through the synchronized queue.
4.3 At Shutdown, Stop Submitting Before You Wait
The safe shutdown order is “stop new submits → wait for completion with WaitForThreadpoolWorkCallbacks → close with CloseThreadpoolWork“. Memory the callback refers to must not be freed before completion is confirmed either. If you free what is referenced while running or pending processing remains, you get use-after-free.6
Thinking “I called the completion wait, so it is safe” is not enough. If another thread can still call SubmitThreadpoolWork, a submit after the wait races with Close. To make the wait a milestone in shutdown, stopping the submitter first is a prerequisite.
If the second argument of WaitForThreadpoolWorkCallbacks is FALSE, it waits for completion; if TRUE, it also requests cancellation of callbacks that have not yet started. That does not mean processing already running may be left behind. How to shut down multiple objects together is covered in Chapter 6.
5. Callback Discipline — Do Not Monopolize or Contaminate a Borrowed Thread
5.1 Declare Long-Running Work, and Check the Return Value Too
The pool adjusts its thread count on the assumption that callbacks return promptly. If you continue long processing or a long wait without telling it anything, the execution of other callbacks is delayed. When the work may take long, either tell the pool about that possibility with CallbackMayRunLong or move it to a dedicated thread.8
However, calling CallbackMayRunLong alone is not enough. It returns FALSE when it cannot make a worker available for other callbacks. Rather than ignoring the return value and continuing to block, in that case lean toward not clogging the pool: split the processing, move it to a dedicated thread, and so on.8
flowchart TB
accTitle: Deciding how to handle a long-running callback
accDescr: Work that needs long processing or a long wait is either moved to a dedicated thread or announced to the pool with CallbackMayRunLong; if FALSE comes back because no other worker could be made available, avoid blocking by splitting the work or moving it to a dedicated thread
long["Needs long processing or waiting"] --> choice{"Where to run it"}
choice -->|"Dedicate"| dedicated["Move to a dedicated thread"]
choice -->|"Run on the pool"| notify["Notify with CallbackMayRunLong"]
notify --> result{"Was another worker secured?"}
result -->|"TRUE"| run["Do the long processing"]
result -->|"FALSE"| split["Split it or move to dedicated"]
Figure 8: Announcing long-running work is a set that includes checking the return value and deciding the next action.
5.2 Do Not Synchronously Wait for Work on the Same Pool from Inside a Worker
A design in which callback A submits job B to the same pool and waits for its completion with WaitForThreadpoolWorkCallbacks or the like needs care. Once every worker is “waiting for another worker’s job”, no free worker remains to run B, and you get a pool-starvation deadlock.
The remedy is not to reserve a worker for waiting but to change to a continuation form in which B’s completion callback submits the next job. Do not write dependencies as waits that occupy a worker.
flowchart TB
accTitle: The structure of a pool-starvation deadlock
accDescr: If every worker thread synchronously waits for the completion of other work submitted to the same pool, no free worker exists that could run that work, and everyone waits forever in a deadlock
w1["Worker 1: waiting for job X to complete"] --> q["Jobs X and Y waiting to run"]
w2["Worker 2: waiting for job Y to complete"] --> q
q -.-> none["No free worker to run them"]
none -.-> dead["Everyone waits forever (starvation deadlock)"]
Figure 9: Synchronously wait for a worker from inside a worker, and no one is left to run the job being waited on.
5.3 Restore the Thread’s State Before Returning
A worker thread is used for the next, unrelated callback as well. Leaving the priority changed, leaving COM initialization state behind, leaving a value in TLS, or forgetting to release a lock carries over into the next job. You must not assume that the function you submit to the pool, or the processing it calls, runs on a dedicated thread.9
There are also APIs that tie cleanup to the end of the callback. For example, LeaveCriticalSectionWhenCallbackReturns asks the pool to release the critical section after the callback returns.4
flowchart TB
accTitle: Do not carry thread state over to the next callback
accDescr: Because a different callback reuses the same worker, returning with priority, COM, TLS, or lock state left behind contaminates the next job; cleaning up and restoring the state prevents that carry-over
first["Callback A borrows a worker"] --> cleanup{"Cleaned up before returning?"}
cleanup -->|"No"| dirty["Reused with state left behind"]
dirty --> impact["Affects unrelated B"]
cleanup -->|"Yes"| clean["Worker returned in its original state"]
clean --> next["Next callback B uses it"]
Figure 10: The thread is borrowed, so you are responsible not only for the result of the processing but for the thread’s state when you return it.
5.4 Do Not Let Unhandled Exceptions Escape the Worker
An unhandled exception on a worker thread can take down the whole process. Just as with a dedicated thread’s thread function, apply the policy of catching exceptions at the callback’s entry point and logging them. Leaving things to the pool does not remove the need to handle failures that occur inside the job.
6. Splitting the Configuration — Custom Pools and Cleanup Groups
6.1 Use a Custom Pool to Isolate Kinds of Work
When the default pool is no longer enough, you can create an independent pool with CreateThreadpool. SetThreadpoolThreadMaximum and SetThreadpoolThreadMinimum set the upper and lower bounds on the number of worker threads.10
The typical purpose is isolation. So that “batch processing that may be slow” does not use up the workers of “work that should respond immediately”, you split the pools and give each a budget of threads. This is not simply about adding threads; it is about separating which work uses which workers.
6.2 Tie the Execution Target and the Cleanup Together with a Callback Environment
What specifies which pool to run on is TP_CALLBACK_ENVIRON. You initialize this callback environment, specify the pool with SetThreadpoolCallbackPool, and pass it to CreateThreadpoolWork and the like. The third argument, which was NULL in the Chapter 4 example, is where the environment goes.5
Through the same environment, a cleanup group can be attached as well. The custom pool is the execution target; the cleanup group is the unit of cleanup. Seeing these two roles separately makes the configuration easier to follow.
flowchart TB
accTitle: Binding the configuration through a callback environment
accDescr: The callback environment points to a custom pool and a cleanup group; work and timer objects created with that environment run on that pool, and the cleanup group's bulk operation gathers the completion wait and release
env["Callback environment (TP_CALLBACK_ENVIRON)"] --> cp["Custom pool (controls the thread count)"]
env --> cg["Cleanup group"]
env --> obj["Passed when creating work / timer / wait / io"]
cg -.-> close["Wait for completion and release in bulk"]
Figure 11: A callback environment is the mechanism that injects “which pool it runs on and who cleans up” at object creation.
6.3 Gather the Completion Wait and Release for Multiple Objects
As work, timer, and other objects multiply within a module, shutdown becomes a list of “wait for each, close each”. If you create a group with CreateThreadpoolCleanupGroup and make the objects you create through the callback environment members of it, a single CloseThreadpoolCleanupGroupMembers gathers the completion wait and release for all member objects.46
Here too, the aim is not to leave a running callback behind. Choose between shutting down work objects individually, as in Chapter 4, and shutting down at the group level, according to the number of objects you manage.
7. Using It from a DLL — Do Not Unload Before the Callbacks
7.1 Wait for Completion in an Explicit Shutdown Function, Not in DllMain
The most dangerous thing in a DLL is for the DLL to be unloaded while its callback code can still execute. If unloaded code runs, you get an access violation.
The basic rule is to stop submitting, wait for the callbacks to complete, and close the objects in the DLL’s explicit shutdown function before unloading. Whether with individual wait functions or with the cleanup group from Chapter 6, this completion check must not be skipped.6
Do not perform this wait inside DllMain. Because of its relationship with the loader lock, it causes a different deadlock. See the article on DllMain and the loader lock for details.
7.2 Pair FreeLibraryWhenCallbackReturns with a Reference Taken Before Submitting
For the situation “this callback is the last job, so I want to release the DLL’s reference after it returns”, there is FreeLibraryWhenCallbackReturns.4
However, this API alone does not prevent an unload before the callback starts. What it does is release one module reference when the running callback returns. Use it as a pair: take a module reference for that processing with GetModuleHandleEx before submitting, and release it from the callback with this API.
flowchart TB
accTitle: Closing the DLL in a shutdown function versus returning a reference from the callback
accDescr: The basic form is a shutdown function other than DllMain that stops submits, waits for completion, and releases before the DLL is unloaded; in a design where the last callback returns the reference, take the reference with GetModuleHandleEx before submitting and pair it with release after return via FreeLibraryWhenCallbackReturns
shutdown["Explicit shutdown function"] --> stop["Stop submits, wait, release"]
stop --> unload["Then unload the DLL"]
note["Do not wait in DllMain"] -.-> shutdown
acquire["Take a module reference before submitting"] --> submit["Submit the callback"]
submit --> callback["Schedule the release inside the callback"]
api["FreeLibraryWhenCallbackReturns"] -.-> callback
callback --> returned["Release one reference after returning"]
Figure 12: Do not confuse the completion wait in the shutdown function with releasing the reference taken for the callback; in both cases, design the DLL’s lifetime first.
8. Summary — Start with work, and Replace Together with Shutdown
The Win32 thread pool is the foundation that changes concurrency in native code from “creating threads” to “handing over jobs as callbacks”. You can consolidate short-lived jobs and waiter-only threads and leave thread management to the OS.
Adoption can be gradual. First, with work, make creation, submission, stopping submits, waiting for completion, and release one set. Next, replace waiter-only threads with wait, and timer-only threads with timer. Consider io integration and pool splitting when they become necessary.
flowchart TB
accTitle: Migrating existing code to the thread pool in stages
accDescr: Review the existing self-managed threads, keep the work that suits the standard library or a dedicated thread, move pool-suited work to work objects including stopping submits and waiting for completion, and replace waits with wait and timers with timer in stages
inventory["Review existing self-managed threads"] --> choose{"Suited to the pool?"}
choose -->|"No"| keep["Choose standard library or dedicated"]
choose -->|"Yes"| work["Move to work with shutdown as a set"]
work --> more["Waits to wait, timers to timer"]
more --> optional["I/O integration, pool splitting as needed"]
Figure 13: Not just reducing threads but completing shutdown at each stage is the basis of a staged migration.
What you keep to the end are three disciplines: do not block for a long time, do not synchronously wait on the same pool, and do not contaminate the thread’s state. In a DLL, additionally prevent races with unloading. Leave what standard C++ or .NET tools can handle to them, and use this API where Windows-specific integration or control is needed.
Related Articles
- The Depths of Windows I/O (Part 3) — I/O Completion Ports (IOCP) and the .NET Thread Pool: The Basement Under async/await
- Practical Multithreading Best Practices: C++ Edition
- Practical Multithreading Best Practices: C Edition
- Spurious Wakeups — Why Condition Variables Wake “Without Being Notified” and How to Wait Correctly on Windows
- DllMain and the Loader Lock — The Real Reason You’re Told to “Do Nothing in DLL Initialization”
- Why You Should Prefer Event Waits over Sleep(1) on Windows
Related Consulting Areas
KomuraSoft LLC handles migration design from native code whose threads have proliferated to the thread pool, design reviews of concurrency in C++ apps and DLLs, and root-cause investigation of hangs and crashes caused by pool starvation or callbacks. You are welcome to consult us starting from an inventory of your existing code.
- Windows Application Development
- Technical Consulting and Design Review
- Bug Investigation and Root-Cause Analysis
- Contact Us
References
-
Microsoft Learn, Thread Pools. On a thread pool being a collection of worker threads that efficiently execute asynchronous callbacks on the app’s behalf; on the types of app it suits (issuing a large number of small work items in parallel, frequently creating and destroying short-lived threads, processing independent work in parallel, exclusive waits on kernel objects, and so on); and on the complete redesign in Vista (unification of worker-thread kinds, a single timer queue, dedicated persistent threads, cleanup groups, multiple pools within a process, and the new API). ↩ ↩2 ↩3 ↩4 ↩5
-
Microsoft Learn, <future>. On task-level asynchronous execution through std::async and future being provided by the standard library, so that concurrency can be written without managing threads directly. ↩
-
Microsoft Learn, Thread Pooling. On the structure of the legacy thread pool API (QueueUserWorkItem, timer queues, registered waits, BindIoCompletionCallback); on there being no way to cancel work once it is queued; and on its explicit statement that the new thread pool API introduced in Vista is simpler and superior in reliability, performance, and flexibility. ↩ ↩2
-
Microsoft Learn, threadpoolapiset.h header. On the function list that includes the four object-creation functions CreateThreadpoolWork, CreateThreadpoolTimer, CreateThreadpoolWait, and CreateThreadpoolIo; cleanup groups (CreateThreadpoolCleanupGroup); and cleanup tied to callback completion (LeaveCriticalSectionWhenCallbackReturns, FreeLibraryWhenCallbackReturns, and so on). ↩ ↩2 ↩3 ↩4
-
Microsoft Learn, CreateThreadpoolWork function (threadpoolapiset.h). On creating a work object from a callback function and a context pointer; and on the third argument, TP_CALLBACK_ENVIRON, being able to specify the callback’s execution environment (the pool it belongs to, and so on), with NULL meaning it runs in the default environment. ↩ ↩2 ↩3
-
Microsoft Learn, Using the Thread Pool Functions. On the basic procedure of creating with CreateThreadpoolWork, submitting with SubmitThreadpoolWork, waiting for completion with WaitForThreadpoolWorkCallbacks, and closing with CloseThreadpoolWork; and on a configuration example that combines a custom pool with a callback environment and a cleanup group. ↩ ↩2 ↩3 ↩4
-
Microsoft Learn, SubmitThreadpoolWork function (threadpoolapiset.h). On being able to submit the same work object multiple times without waiting for a preceding callback to complete, so that callbacks run in parallel; and on the pool being able to adjust (throttle) the number of threads for efficiency. ↩
-
Microsoft Learn, CallbackMayRunLong function (threadpoolapiset.h). On notifying the pool that the current callback may run for a long time, which the pool uses to decide whether to secure threads for other callbacks; and on considering a dedicated thread for a long-running callback where possible. ↩ ↩2
-
Microsoft Learn, Thread Pooling. On work items submitted to the thread pool, and the functions they call, having to be thread-pool safe; on not assuming that the executing thread is a dedicated, persistent thread; and on avoiding the use of TLS and asynchronous calls that require a persistent thread. ↩
-
Microsoft Learn, SetThreadpoolThreadMaximum function (threadpoolapiset.h). On being able to set an upper bound on the number of worker threads for a pool created with CreateThreadpool (the lower bound is SetThreadpoolThreadMinimum). ↩
Related Articles
Recent articles sharing the same tags. Deepen your understanding with closely related topics.
DllMain and the Loader Lock — The Real Reason You Are Told to "Do Nothing in DLL Initialization"
Why DllMain must not call LoadLibrary or wait on threads: the loader lock serializes DLL notifications, typical deadlocks, deferred initi...
Why Arguments Break — The Rules of Windows Command-Line Arguments
Windows passes CreateProcess a single string that the receiver splits. Covers the CommandLineToArgvW, CRT, and .NET rules, ArgumentList, ...
What Is Left After the Parent Dies — Keeping Child Processes in a Job Object
Why do SDK helpers outlive a killed UI and keep the camera or COM port? Designing child process lifetime with Job Objects, KillOnJobClose...
Named Pipes in Practice — Windows' Standard IPC from Design to Security
A practical guide to named pipes, Windows' standard IPC. Covers byte vs. message mode, servers that handle multiple clients, ACL and impe...
What "Not Responding" Really Is — How Windows Decides an App Has Hung, and How to Design Apps That Don't
Windows marks a window Not Responding after 5 seconds without message retrieval and shows a ghost window: the check, hang causes, UI-thre...
Related Topics
These topic pages place the article in a broader service and decision context.
Windows Technical Topics
Topic hub for KomuraSoft LLC's Windows development, investigation, and legacy-asset articles.
Where This Topic Connects
This article connects naturally to the following service pages.
Windows App Development
We support Windows desktop applications that involve resident processing, device integration, operational logging, and maintainable structure.
Frequently Asked Questions
Common questions about the topic of this article.
- Compared with creating my own threads with CreateThread, what is better about a thread pool?
- Efficiency when you process a large number of short-lived jobs, and less thread-management code. Creating and destroying a thread has a cost you cannot ignore, so an app that repeats "CreateThread for every job and destroy it when done", or that holds many threads that sleep only to wait for an event, can reduce its thread count and context switches by moving to a pool. The official documentation also lists as pool candidates apps that issue a large number of small work items in parallel, apps that create many short-lived threads, and apps that have threads dedicated to waiting on kernel objects. Conversely, work in which "the thread itself needs a personality", such as a priority change, a COM STA, or dedicated processing that keeps running for a long time, should stay on a dedicated thread as before.
- How does it differ from the older thread pool functions such as QueueUserWorkItem?
- The thread pool was completely redesigned in Windows Vista. The current threadpoolapiset family (CreateThreadpoolWork and so on) is the new API; QueueUserWorkItem, RegisterWaitForSingleObject, and the like are the old (legacy) API. The new API unifies the kinds of worker thread, lets you create multiple independent pools within a process, and provides mechanisms such as bulk release through a cleanup group and lock release or DLL unload tied to callback completion. The official documentation also states that the new API is simpler and superior in reliability, performance, and flexibility. The legacy API has structural constraints too, such as "there is no way to cancel work once it is queued", so use the new API in new code.
- Are there things I must not do inside a callback?
- There are three main ones. First, blocking for a long time or running long processing at the defaults. The pool adjusts its thread count on the assumption that callbacks finish promptly, so for processing that takes long, either declare it with CallbackMayRunLong or use a dedicated thread. Second, synchronously waiting for the completion of other work submitted to the same pool. When every worker ends up "waiting for another worker", you get a pool-starvation deadlock. Third, depending on the thread's personality. Worker threads are shared between callbacks, so returning with the thread priority or COM initialization state changed, or leaving state in TLS, contaminates the next callback. For cleanup at the end (releasing a lock or unloading a DLL), dedicated mechanisms such as LeaveCriticalSectionWhenCallbackReturns and FreeLibraryWhenCallbackReturns are provided.
- Is there anything to watch out for when using the thread pool inside a DLL?
- The biggest danger is "the DLL is unloaded while a callback is still running". If a callback executes after the unload, you get an access violation. In its shutdown processing, the DLL must reliably wait for the callbacks it issued to complete, using a wait function such as WaitForThreadpoolWorkCallbacks (or CloseThreadpoolCleanupGroupMembers on a cleanup group), before closing the objects. However, doing this wait inside DllMain can deadlock through its interaction with the loader lock, so the rule is to do it in an explicit shutdown function, not in DllMain. For the case where the callback itself is "the last job" and wants to free the DLL, a dedicated API, FreeLibraryWhenCallbackReturns, is provided.
- Now that C++ has std::async and .NET has ThreadPool, are there still occasions to use this API directly?
- Yes. The criterion is "whether the tool at that layer is enough". If the granularity of concurrency you need is covered by std::async or std::thread in C++, the standard library is the first candidate, from a portability standpoint as well. On the other hand, wanting to unify timers, kernel object waits, and asynchronous I/O completion in one callback mechanism, wanting to split pools and control the thread count per kind of work, or not wanting to own threads inside a DLL or COM component are requirements the Win32 thread pool covers. The relationship with .NET's ThreadPool and IOCP is covered in a related article, and as long as you write native code, knowing this mechanism in the layer below is never wasted.