The Win32 Thread Pool API — Concurrency That Does Not Create Threads, with CreateThreadpoolWork

· Updated: · · Windows, Multithreading, C++, Windows Development, Win32 API, Performance Improvement

Revision history (1 updates, last updated Sep 7, 2026)

A log of the changes made to this article. Where a pre-update version was archived, it stays readable at a permanent DOI link.

Retranslated as a full translation of the current Japanese original. The previous English version was an abridgement that dropped subsections, tables, diagrams, and paragraphs; all of them have been restored to match the Japanese article, and the knowledge map section has been added where the Japanese article has one. The technical claims are the same as in the Japanese version. Read the version before this update (DOI: 10.5281/zenodo.22170915)
First published
Cite this article(DOI: 10.5281/zenodo.22170914)

This article is archived on Zenodo. Below are both the DOI that always resolves to the latest version and the DOI pinned to the version you are reading.

Go Komura (2026). The Win32 Thread Pool API — Concurrency That Does Not Create Threads, with CreateThreadpoolWork. KomuraSoft LLC. https://doi.org/10.5281/zenodo.22170914 https://comcomponent.com/en/blog/win32-thread-pool-api/

DOI (latest version)
10.5281/zenodo.22170914
DOI (this version)
10.5281/zenodo.22615677

“One CreateThread per client.” “One thread for the timer.” “One for waiting on an event.” In native Windows code, the reflex is to add a thread for every small job. Each thread consumes a stack and a kernel object, and creating and destroying it costs something as well.

What absorbs this mismatch is the OS-standard Win32 thread pool API. The app hands over the processing it wants to run as a callback and leaves worker-thread management to the OS. “Not creating threads” does not mean that threads become unnecessary; it means that the app does not create and destroy one itself for every job.1

This article is for developers who write Windows apps, services, and DLLs in C/C++, and proceeds in the order choosing the tool → selecting the object → implementing work → cautions for callbacks and shutdown. The subject is the CreateThreadpoolWork family of APIs redesigned in Windows Vista. The code is an excerpt that shows the skeleton of usage; app-specific parts such as the queue are omitted.

1. The Bottom Line First

If you process a large number of short-lived jobs, manage jobs rather than threads. The app, however, still designs when a job stops and when it can be released.

There are three axes of judgment.

  • Choose the tool. If standard C++ or .NET tools are enough, use them. This API’s turn comes when you want to unify timers, waits, and I/O completion in Win32, or when you want to split pools.
  • Hand over work in units of jobs. Use the new API’s work, timer, wait, and io, and leave jobs that need thread-specific state, such as a thread priority or a COM STA, on dedicated threads.
  • Build shutdown as part of the set. The basic form is “stop submitting → wait for completion → close”. Inside a callback, do not block for a long time, do not synchronously wait for work on the same pool, and restore the thread’s state.

If you want to get something running first, start with the work example in Chapter 4 and then check the callback discipline in Chapter 5. Bulk shutdown of multiple objects is covered in Chapter 6, and DLL unloading in Chapter 7.

2. Choosing Between a Pool, a Dedicated Thread, and the Standard Library

2.1 A Pool Suits Short-Lived, High-Volume, Wait-Centric Work

A thread pool is a set of worker threads managed by the OS. The workers execute callbacks one after another, and the OS adjusts their number according to load. By handing over the work you used to do by creating and tearing down your own threads, you reduce both management code and the cost of creation and destruction.1

The official documentation lists as candidates apps that issue a large number of small work items in parallel, apps that frequently create and destroy short-lived threads, apps that process independent work in parallel in the background, and apps that have dedicated threads waiting on kernel objects or events. Search, network I/O, and consolidating waiter-only threads are typical.1

On the other hand, work that needs a change of thread priority, needs a COM STA, or keeps running for the whole lifetime of the process stays on a dedicated thread. The criterion is whether it is enough for the processing to run, or whether the thread itself needs a “personality”.

Choosing between a dedicated thread and a poolFirst check whether the thread needs a personality such as a priority or an STA and whether it runs for a long time; only short-lived, high-volume, or wait-style work that matches neither goes on the thread poolYesNoYesNoNeeds a personality such as priority or STA?Keep on a dedicated threadRuns for a long time?Put on the thread poolShort-lived work, waits, timers, I/O completions

Figure 1: Only work that “needs no personality and finishes quickly” belongs on the pool. Everything else stays on a dedicated thread as before.

2.2 First Ask Whether Standard C++ or .NET Is Enough

If the granularity is covered by C++ std::async or std::thread, the standard library is the first candidate. It is portable, and the behavior of std::async and future can be handled on the basis of the standard.2

The reasons to use the Win32 thread pool directly are requirements such as wanting a unified callback mechanism that includes timer, wait, and io; wanting separate pools or thread counts per kind of work; or not wanting to own threads inside a DLL or COM component. Separate the question of whether you simply want parallelism from whether you need Windows-specific control.

Deciding which layer's tools to write withIf standard C++ async or thread is enough, use it; use the Win32 thread pool directly when you need to unify timers, waits, and I/O completion, split pools or control thread counts, or avoid owning threads inside a DLL or COM componentYesNoUnify timer, wait, and ioSplit pools / control countsAvoid own threads in a DLLAre the standard C++ tools enough?std::async / std::threadWhat do you need?Win32 thread pool

Figure 2: When in doubt, start with the standard library; this API’s turn comes when a requirement it cannot express appears.

In .NET, ThreadPool and Task play the same role, and I/O completion is tied to IOCP. For a detailed explanation of the relationship, see the article on IOCP and the .NET thread pool. There is no need to go down to the Win32 API where the tools in the layer above are enough.

2.3 In New Code, Use the New API from Vista Onward

There are two generations of Win32 thread pool API: the legacy API that has continued since Windows 2000, such as QueueUserWorkItem and RegisterWaitForSingleObject, and the new CreateThreadpoolWork family that was completely redesigned in Windows Vista.

The new API unified the kinds of worker thread and provided a single timer queue, dedicated persistent threads, multiple independent pools within a process, and cleanup groups. The official documentation also cites the new API’s simplicity, reliability, performance, and flexibility as its advantages.13

The legacy API also has a structural constraint: there is no way to cancel work once it has been queued. Use the new API in new code, and when reviewing existing code, start from the following correspondence.3

Correspondence between the legacy thread pool API and the new APIThe legacy API's QueueUserWorkItem is replaced by the new API's work object, timer queues by timer, registered waits by wait, and BindIoCompletionCallback by ioQueueUserWorkItemworkTimer queuestimerRegistered waitswaitBindIoCompletionCallbackio

Figure 3: The migration target from the legacy API is fixed one-to-one. An inventory of existing code can start from this correspondence table.

3. How It Works — Four Objects with Different Firing Conditions

3.1 Choose by What Should Trigger the Callback

At the center of the new API are the following four kinds. Choose not only by “what to process” but by what should trigger the callback.4

Object Creation function Callback firing condition
work CreateThreadpoolWork When submitted with SubmitThreadpoolWork
timer CreateThreadpoolTimer When the specified time or period arrives
wait CreateThreadpoolWait When a kernel object becomes signaled
io CreateThreadpoolIo When asynchronous I/O on the associated handle completes

The firing conditions differ, but the workers of the same pool do the executing. Instead of writing periodic processing, event response, and I/O completion handling each on its own dedicated thread, you can align them on a common callback mechanism.

Separating firing conditions from callback executionThe pool's workers execute callbacks that objects with firing conditions have made ready, and a worker that has finished is used for the next callback as well, so the app does not create a thread per jobApp specifies the job and its firing conditionObject waits for the conditionCallback becomes ready to runPool workers execute itFinish and return the workerReused for the next callback too

Figure 4: Separating the mechanism that waits for a trigger from the workers that execute the processing reduces per-job thread management.

3.2 “Threads That Only Sleep and Wait” Can Be Replaced Too

The pool’s benefit is not limited to CPU-bound work. Timers are consolidated into a single timer queue, and waits on multiple handles onto a small number of wait threads.1

For example, if you have five threads that sleep only so they can “run when the event is signaled”, you can think of replacing them with five wait objects. Instead of holding a thread that sleeps on its own for each one, you run the on-signal processing as a callback.

Replacing waiter-only threads with wait objectsWaiter-only threads that slept one per event are consolidated onto the pool's wait threads when replaced with wait objects, and the callback runs only when signaled5 waiter-only threads sleeping individuallyConsumes 5 stacks and 5 threads5 wait objectsConsolidated onto the pool's wait threadsCallback runs only on signal

Figure 5: “Threads that only sleep and wait” can be eliminated by turning them into wait objects. This is the obvious first move in a pool migration.

4. Implementation Basics — Create a work, Submit It, and Shut Down Safely

4.1 First, One Round Trip from Creation to Shutdown

Walk through usage once with the most basic object, work. CreateThreadpoolWork binds the callback to a context, and SubmitThreadpoolWork requests execution. If the third argument is NULL, the process’s default pool is used. Many uses are fine with this default pool.56

The following is an excerpt that shows the procedure, not a complete program that compiles as is. WORK_QUEUE, ITEM, Enqueue, Dequeue, and ProcessItem stand for app-side processing. Implement the queue’s mutual exclusion, stopping the submitter, and failure handling separately, and if creating the work fails, do not proceed to the subsequent submit, wait, and release.

VOID CALLBACK WorkCallback(PTP_CALLBACK_INSTANCE instance,
                           PVOID context, PTP_WORK work)
{
    // The context is fixed at creation time. Per-item data is passed through a synchronized queue
    WORK_QUEUE* queue = (WORK_QUEUE*)context;
    ITEM* item = Dequeue(queue);        // Take one item under mutual exclusion
    ProcessItem(item);
}

// 1) Create (bind the callback to the shared context, i.e. the queue)
PTP_WORK work = CreateThreadpoolWork(WorkCallback, &queue, NULL);
if (!work) { /* Failure handling with GetLastError */ }

// 2) Submit once for each item you enqueue (keep the item count and the submit count equal)
Enqueue(&queue, item);
SubmitThreadpoolWork(work);

// 3) Stop the submitter, then wait for completion (TRUE also attempts to cancel work that has not started)
WaitForThreadpoolWorkCallbacks(work, FALSE);

// 4) Close
CloseThreadpoolWork(work);
Lifecycle of a work objectCreate with CreateThreadpoolWork; submitting with SubmitThreadpoolWork runs callbacks in parallel. At shutdown, first stop new submits, wait for all callbacks to complete with WaitForThreadpoolWorkCallbacks, then close with CloseThreadpoolWorkCreate with CreateThreadpoolWorkSubmit with SubmitThreadpoolWork (repeatable)Callbacks run in parallelStop new submitsWait for completion with WaitForThreadpoolWorkCallbacksClose with CloseThreadpoolWork

Figure 6: The shutdown order is “stop submitting → wait for completion → close”. Skip any step and you get use-after-free or a race.

4.2 A work Can Be Reused, but the context Does Not Change per Submit

The same work object can be submitted multiple times, even before the previous callback has finished. Each submit runs the callback, and multiple instances run in parallel. The number of threads actually used may be adjusted by the pool for efficiency.7

What to distinguish here is the work from the data of each work item. The context passed to the callback is fixed at creation time. To process N items of different data with one work, make a synchronized queue the context, as in the example above. Submit once for each item you enqueue, and have the callback take one item from the queue. A design that creates a work object per item is also fine.5

Receiving per-item data through a fixed contextThe context is fixed to a shared queue when the work is created; the submitter submits once for each item it enqueues, and each callback running in parallel takes one item at a time from the queue under mutual exclusionFix the context at work creationShared queue with mutual exclusionEnqueue one itemSubmit once per itemCallbacks run in parallelDequeue one item eachProcess the item taken

Figure 7: What is fixed is the context that points to the queue; per-item data is passed through the synchronized queue.

4.3 At Shutdown, Stop Submitting Before You Wait

The safe shutdown order is “stop new submits → wait for completion with WaitForThreadpoolWorkCallbacks → close with CloseThreadpoolWork“. Memory the callback refers to must not be freed before completion is confirmed either. If you free what is referenced while running or pending processing remains, you get use-after-free.6

Thinking “I called the completion wait, so it is safe” is not enough. If another thread can still call SubmitThreadpoolWork, a submit after the wait races with Close. To make the wait a milestone in shutdown, stopping the submitter first is a prerequisite.

If the second argument of WaitForThreadpoolWorkCallbacks is FALSE, it waits for completion; if TRUE, it also requests cancellation of callbacks that have not yet started. That does not mean processing already running may be left behind. How to shut down multiple objects together is covered in Chapter 6.

5. Callback Discipline — Do Not Monopolize or Contaminate a Borrowed Thread

5.1 Declare Long-Running Work, and Check the Return Value Too

The pool adjusts its thread count on the assumption that callbacks return promptly. If you continue long processing or a long wait without telling it anything, the execution of other callbacks is delayed. When the work may take long, either tell the pool about that possibility with CallbackMayRunLong or move it to a dedicated thread.8

However, calling CallbackMayRunLong alone is not enough. It returns FALSE when it cannot make a worker available for other callbacks. Rather than ignoring the return value and continuing to block, in that case lean toward not clogging the pool: split the processing, move it to a dedicated thread, and so on.8

Deciding how to handle a long-running callbackWork that needs long processing or a long wait is either moved to a dedicated thread or announced to the pool with CallbackMayRunLong; if FALSE comes back because no other worker could be made available, avoid blocking by splitting the work or moving it to a dedicated threadDedicateRun on the poolTRUEFALSENeeds long processing or waitingWhere to run itMove to a dedicated threadNotify with CallbackMayRunLongWas another worker secured?Do the long processingSplit it or move to dedicated

Figure 8: Announcing long-running work is a set that includes checking the return value and deciding the next action.

5.2 Do Not Synchronously Wait for Work on the Same Pool from Inside a Worker

A design in which callback A submits job B to the same pool and waits for its completion with WaitForThreadpoolWorkCallbacks or the like needs care. Once every worker is “waiting for another worker’s job”, no free worker remains to run B, and you get a pool-starvation deadlock.

The remedy is not to reserve a worker for waiting but to change to a continuation form in which B’s completion callback submits the next job. Do not write dependencies as waits that occupy a worker.

The structure of a pool-starvation deadlockIf every worker thread synchronously waits for the completion of other work submitted to the same pool, no free worker exists that could run that work, and everyone waits forever in a deadlockWorker 1: waiting for job X to completeJobs X and Y waiting to runWorker 2: waiting for job Y to completeNo free worker to run themEveryone waits forever (starvation deadlock)

Figure 9: Synchronously wait for a worker from inside a worker, and no one is left to run the job being waited on.

5.3 Restore the Thread’s State Before Returning

A worker thread is used for the next, unrelated callback as well. Leaving the priority changed, leaving COM initialization state behind, leaving a value in TLS, or forgetting to release a lock carries over into the next job. You must not assume that the function you submit to the pool, or the processing it calls, runs on a dedicated thread.9

There are also APIs that tie cleanup to the end of the callback. For example, LeaveCriticalSectionWhenCallbackReturns asks the pool to release the critical section after the callback returns.4

Do not carry thread state over to the next callbackBecause a different callback reuses the same worker, returning with priority, COM, TLS, or lock state left behind contaminates the next job; cleaning up and restoring the state prevents that carry-overNoYesCallback A borrows a workerCleaned up before returning?Reused with state left behindAffects unrelated BWorker returned in its original stateNext callback B uses it

Figure 10: The thread is borrowed, so you are responsible not only for the result of the processing but for the thread’s state when you return it.

5.4 Do Not Let Unhandled Exceptions Escape the Worker

An unhandled exception on a worker thread can take down the whole process. Just as with a dedicated thread’s thread function, apply the policy of catching exceptions at the callback’s entry point and logging them. Leaving things to the pool does not remove the need to handle failures that occur inside the job.

6. Splitting the Configuration — Custom Pools and Cleanup Groups

6.1 Use a Custom Pool to Isolate Kinds of Work

When the default pool is no longer enough, you can create an independent pool with CreateThreadpool. SetThreadpoolThreadMaximum and SetThreadpoolThreadMinimum set the upper and lower bounds on the number of worker threads.10

The typical purpose is isolation. So that “batch processing that may be slow” does not use up the workers of “work that should respond immediately”, you split the pools and give each a budget of threads. This is not simply about adding threads; it is about separating which work uses which workers.

6.2 Tie the Execution Target and the Cleanup Together with a Callback Environment

What specifies which pool to run on is TP_CALLBACK_ENVIRON. You initialize this callback environment, specify the pool with SetThreadpoolCallbackPool, and pass it to CreateThreadpoolWork and the like. The third argument, which was NULL in the Chapter 4 example, is where the environment goes.5

Through the same environment, a cleanup group can be attached as well. The custom pool is the execution target; the cleanup group is the unit of cleanup. Seeing these two roles separately makes the configuration easier to follow.

Binding the configuration through a callback environmentThe callback environment points to a custom pool and a cleanup group; work and timer objects created with that environment run on that pool, and the cleanup group's bulk operation gathers the completion wait and releaseCallback environment (TP_CALLBACK_ENVIRON)Custom pool (controls the thread count)Cleanup groupPassed when creating work / timer / wait / ioWait for completion and release in bulk

Figure 11: A callback environment is the mechanism that injects “which pool it runs on and who cleans up” at object creation.

6.3 Gather the Completion Wait and Release for Multiple Objects

As work, timer, and other objects multiply within a module, shutdown becomes a list of “wait for each, close each”. If you create a group with CreateThreadpoolCleanupGroup and make the objects you create through the callback environment members of it, a single CloseThreadpoolCleanupGroupMembers gathers the completion wait and release for all member objects.46

Here too, the aim is not to leave a running callback behind. Choose between shutting down work objects individually, as in Chapter 4, and shutting down at the group level, according to the number of objects you manage.

7. Using It from a DLL — Do Not Unload Before the Callbacks

7.1 Wait for Completion in an Explicit Shutdown Function, Not in DllMain

The most dangerous thing in a DLL is for the DLL to be unloaded while its callback code can still execute. If unloaded code runs, you get an access violation.

The basic rule is to stop submitting, wait for the callbacks to complete, and close the objects in the DLL’s explicit shutdown function before unloading. Whether with individual wait functions or with the cleanup group from Chapter 6, this completion check must not be skipped.6

Do not perform this wait inside DllMain. Because of its relationship with the loader lock, it causes a different deadlock. See the article on DllMain and the loader lock for details.

7.2 Pair FreeLibraryWhenCallbackReturns with a Reference Taken Before Submitting

For the situation “this callback is the last job, so I want to release the DLL’s reference after it returns”, there is FreeLibraryWhenCallbackReturns.4

However, this API alone does not prevent an unload before the callback starts. What it does is release one module reference when the running callback returns. Use it as a pair: take a module reference for that processing with GetModuleHandleEx before submitting, and release it from the callback with this API.

Closing the DLL in a shutdown function versus returning a reference from the callbackThe basic form is a shutdown function other than DllMain that stops submits, waits for completion, and releases before the DLL is unloaded; in a design where the last callback returns the reference, take the reference with GetModuleHandleEx before submitting and pair it with release after return via FreeLibraryWhenCallbackReturnsExplicit shutdown functionStop submits, wait, releaseThen unload the DLLDo not wait in DllMainTake a module reference before submittingSubmit the callbackSchedule the release inside the callbackFreeLibraryWhenCallbackReturnsRelease one reference after returning

Figure 12: Do not confuse the completion wait in the shutdown function with releasing the reference taken for the callback; in both cases, design the DLL’s lifetime first.

8. Summary — Start with work, and Replace Together with Shutdown

The Win32 thread pool is the foundation that changes concurrency in native code from “creating threads” to “handing over jobs as callbacks”. You can consolidate short-lived jobs and waiter-only threads and leave thread management to the OS.

Adoption can be gradual. First, with work, make creation, submission, stopping submits, waiting for completion, and release one set. Next, replace waiter-only threads with wait, and timer-only threads with timer. Consider io integration and pool splitting when they become necessary.

Migrating existing code to the thread pool in stagesReview the existing self-managed threads, keep the work that suits the standard library or a dedicated thread, move pool-suited work to work objects including stopping submits and waiting for completion, and replace waits with wait and timers with timer in stagesNoYesReview existing self-managed threadsSuited to the pool?Choose standard library or dedicatedMove to work with shutdown as a setWaits to wait, timers to timerI/O integration, pool splitting as needed

Figure 13: Not just reducing threads but completing shutdown at each stage is the basis of a staged migration.

What you keep to the end are three disciplines: do not block for a long time, do not synchronously wait on the same pool, and do not contaminate the thread’s state. In a DLL, additionally prevent races with unloading. Leave what standard C++ or .NET tools can handle to them, and use this API where Windows-specific integration or control is needed.

KomuraSoft LLC handles migration design from native code whose threads have proliferated to the thread pool, design reviews of concurrency in C++ apps and DLLs, and root-cause investigation of hangs and crashes caused by pool starvation or callbacks. You are welcome to consult us starting from an inventory of your existing code.

References

  1. Microsoft Learn, Thread Pools. On a thread pool being a collection of worker threads that efficiently execute asynchronous callbacks on the app’s behalf; on the types of app it suits (issuing a large number of small work items in parallel, frequently creating and destroying short-lived threads, processing independent work in parallel, exclusive waits on kernel objects, and so on); and on the complete redesign in Vista (unification of worker-thread kinds, a single timer queue, dedicated persistent threads, cleanup groups, multiple pools within a process, and the new API). ↩ ↩2 ↩3 ↩4 ↩5

  2. Microsoft Learn, <future>. On task-level asynchronous execution through std::async and future being provided by the standard library, so that concurrency can be written without managing threads directly. ↩

  3. Microsoft Learn, Thread Pooling. On the structure of the legacy thread pool API (QueueUserWorkItem, timer queues, registered waits, BindIoCompletionCallback); on there being no way to cancel work once it is queued; and on its explicit statement that the new thread pool API introduced in Vista is simpler and superior in reliability, performance, and flexibility. ↩ ↩2

  4. Microsoft Learn, threadpoolapiset.h header. On the function list that includes the four object-creation functions CreateThreadpoolWork, CreateThreadpoolTimer, CreateThreadpoolWait, and CreateThreadpoolIo; cleanup groups (CreateThreadpoolCleanupGroup); and cleanup tied to callback completion (LeaveCriticalSectionWhenCallbackReturns, FreeLibraryWhenCallbackReturns, and so on). ↩ ↩2 ↩3 ↩4

  5. Microsoft Learn, CreateThreadpoolWork function (threadpoolapiset.h). On creating a work object from a callback function and a context pointer; and on the third argument, TP_CALLBACK_ENVIRON, being able to specify the callback’s execution environment (the pool it belongs to, and so on), with NULL meaning it runs in the default environment. ↩ ↩2 ↩3

  6. Microsoft Learn, Using the Thread Pool Functions. On the basic procedure of creating with CreateThreadpoolWork, submitting with SubmitThreadpoolWork, waiting for completion with WaitForThreadpoolWorkCallbacks, and closing with CloseThreadpoolWork; and on a configuration example that combines a custom pool with a callback environment and a cleanup group. ↩ ↩2 ↩3 ↩4

  7. Microsoft Learn, SubmitThreadpoolWork function (threadpoolapiset.h). On being able to submit the same work object multiple times without waiting for a preceding callback to complete, so that callbacks run in parallel; and on the pool being able to adjust (throttle) the number of threads for efficiency. ↩

  8. Microsoft Learn, CallbackMayRunLong function (threadpoolapiset.h). On notifying the pool that the current callback may run for a long time, which the pool uses to decide whether to secure threads for other callbacks; and on considering a dedicated thread for a long-running callback where possible. ↩ ↩2

  9. Microsoft Learn, Thread Pooling. On work items submitted to the thread pool, and the functions they call, having to be thread-pool safe; on not assuming that the executing thread is a dedicated, persistent thread; and on avoiding the use of TLS and asynchronous calls that require a persistent thread. ↩

  10. Microsoft Learn, SetThreadpoolThreadMaximum function (threadpoolapiset.h). On being able to set an upper bound on the number of worker threads for a pool created with CreateThreadpool (the lower bound is SetThreadpoolThreadMinimum). ↩

Recent articles sharing the same tags. Deepen your understanding with closely related topics.

These topic pages place the article in a broader service and decision context.

This article connects naturally to the following service pages.

Frequently Asked Questions

Common questions about the topic of this article.

Compared with creating my own threads with CreateThread, what is better about a thread pool?
Efficiency when you process a large number of short-lived jobs, and less thread-management code. Creating and destroying a thread has a cost you cannot ignore, so an app that repeats "CreateThread for every job and destroy it when done", or that holds many threads that sleep only to wait for an event, can reduce its thread count and context switches by moving to a pool. The official documentation also lists as pool candidates apps that issue a large number of small work items in parallel, apps that create many short-lived threads, and apps that have threads dedicated to waiting on kernel objects. Conversely, work in which "the thread itself needs a personality", such as a priority change, a COM STA, or dedicated processing that keeps running for a long time, should stay on a dedicated thread as before.
How does it differ from the older thread pool functions such as QueueUserWorkItem?
The thread pool was completely redesigned in Windows Vista. The current threadpoolapiset family (CreateThreadpoolWork and so on) is the new API; QueueUserWorkItem, RegisterWaitForSingleObject, and the like are the old (legacy) API. The new API unifies the kinds of worker thread, lets you create multiple independent pools within a process, and provides mechanisms such as bulk release through a cleanup group and lock release or DLL unload tied to callback completion. The official documentation also states that the new API is simpler and superior in reliability, performance, and flexibility. The legacy API has structural constraints too, such as "there is no way to cancel work once it is queued", so use the new API in new code.
Are there things I must not do inside a callback?
There are three main ones. First, blocking for a long time or running long processing at the defaults. The pool adjusts its thread count on the assumption that callbacks finish promptly, so for processing that takes long, either declare it with CallbackMayRunLong or use a dedicated thread. Second, synchronously waiting for the completion of other work submitted to the same pool. When every worker ends up "waiting for another worker", you get a pool-starvation deadlock. Third, depending on the thread's personality. Worker threads are shared between callbacks, so returning with the thread priority or COM initialization state changed, or leaving state in TLS, contaminates the next callback. For cleanup at the end (releasing a lock or unloading a DLL), dedicated mechanisms such as LeaveCriticalSectionWhenCallbackReturns and FreeLibraryWhenCallbackReturns are provided.
Is there anything to watch out for when using the thread pool inside a DLL?
The biggest danger is "the DLL is unloaded while a callback is still running". If a callback executes after the unload, you get an access violation. In its shutdown processing, the DLL must reliably wait for the callbacks it issued to complete, using a wait function such as WaitForThreadpoolWorkCallbacks (or CloseThreadpoolCleanupGroupMembers on a cleanup group), before closing the objects. However, doing this wait inside DllMain can deadlock through its interaction with the loader lock, so the rule is to do it in an explicit shutdown function, not in DllMain. For the case where the callback itself is "the last job" and wants to free the DLL, a dedicated API, FreeLibraryWhenCallbackReturns, is provided.
Now that C++ has std::async and .NET has ThreadPool, are there still occasions to use this API directly?
Yes. The criterion is "whether the tool at that layer is enough". If the granularity of concurrency you need is covered by std::async or std::thread in C++, the standard library is the first candidate, from a portability standpoint as well. On the other hand, wanting to unify timers, kernel object waits, and asynchronous I/O completion in one callback mechanism, wanting to split pools and control the thread count per kind of work, or not wanting to own threads inside a DLL or COM component are requirements the Win32 thread pool covers. The relationship with .NET's ThreadPool and IOCP is covered in a related article, and as long as you write native code, knowing this mechanism in the layer below is never wasted.

Author Profile

Profile page for the article author.

Go Komura

Representative of KomuraSoft LLC

Focused on Windows software development, technical consulting, and investigations into failures that are difficult to reproduce.

Back to the Blog