How to Convert YUV to RGB with Media Foundation
· Updated: · Go Komura · Media Foundation, C++, Windows Development, Video Processing, YUV
Revision history (1 updates, last updated Sep 1, 2026)
A log of the changes made to this article. Where a pre-update version was archived, it stays readable at a permanent DOI link.
- Retranslated as a full translation of the Japanese original. The previous English version was an abridgement that carried only part of the source, so sections, tables, Mermaid diagrams, figure captions and FAQ entries were missing. All of them have been restored to match the Japanese original, and the technical claims are the same as in the Japanese version. Read the version before this update (DOI: 10.5281/zenodo.21614501)
- First published
Cite this article(DOI: 10.5281/zenodo.21614500)
This article is archived on Zenodo. Below are both the DOI that always resolves to the latest version and the DOI pinned to the version you are reading.
Go Komura (2026). How to Convert YUV to RGB with Media Foundation. KomuraSoft LLC. https://doi.org/10.5281/zenodo.21614500 https://comcomponent.com/en/blog/2026/03/15/002-media-foundation-yuv-to-rgb-conversion-patterns/
- DOI (latest version)
- 10.5281/zenodo.21614500
- DOI (this version)
- 10.5281/zenodo.22217146
You want to pull a frame out of a video and save it as a PNG, pass it to WIC or GDI, or display it in your UI. In those situations, the application side wants RGB pixel data.
However, the frames that come out of a Media Foundation decoder are, quite commonly, YUV-family formats such as NV12 or YUY2. If you treat the raw byte stream as an image as-is, you get a rather sad picture: broken colors, banding, or a strangely greenish tint.
In an earlier post, An Introduction to Media Foundation - Understanding the API Through a COM Lens, we covered the big picture, and in Extracting a Still Image from an MP4 at a Specific Time with Media Foundation, we covered still-image extraction. This time we tackle the step that sits in between: the YUV -> RGB conversion itself.
In this article, we separate and organize the following two patterns.
- Pattern A: let
IMFSourceReaderautomatically take the frames all the way to RGB32 - Pattern B: receive
NV12/YUY2and convert to RGB yourself
The goal is not to memorize API names. It is to be able to picture, in your head, where in Media Foundation the YUV appears and where it turns into RGB.
The code that appears in this article is published on GitHub as a complete sample set (C++ code for Pattern A / Pattern B, CMake configuration, and tests for the pixel conversion).
media-foundation-yuv-to-rgb-conversion-patterns - komurasoft-blog-samples (GitHub)
Prerequisites for Building and Running
If you are going to carry the code in this article into your own project, this is all you need.
| Item | Requirement |
|---|---|
| OS | Windows 10 or later |
| Compiler | MSVC from Visual Studio 2019 / 2022 (C++17) |
| SDK | Windows SDK (the Media Foundation headers and import libraries). It is included in the Desktop development with C++ workload in Visual Studio |
| Build | The sample uses CMake 3.20 or later. Creating a Visual Studio project by hand is fine too |
There are 4 libraries to link. The article code writes them with #pragma comment(lib, ...), but specifying them in the project settings works exactly the same.
mfplat.libmfreadwrite.libmfuuid.libole32.lib
Also, the code in this article assumes CoInitializeEx and MFStartup have already been called. Only the single-pixel conversion formula (5.6.) is OS-independent, so the GitHub sample factors it out into a separate header that can also be tested with g++ on Linux.
1. The Conclusion First
Summarizing the conclusions up front:
- For extracting a few still images or generating thumbnails, the easiest route is to enable
MF_SOURCE_READER_ENABLE_VIDEO_PROCESSINGand requestMFVideoFormat_RGB32 - However, this automatic conversion is software processing and is not optimized for real-time playback
- If you are going to write your own conversion, the shortest path is to properly understand
NV12andYUY2first - YUV -> RGB is not just “multiply by three coefficients and you are done” — in practice, subsampling, range, matrix, and stride all come into play
- The Media Foundation documentation broadly uses the term
YUV, but for digital video it is easier to read if you assume it effectively means Y’CbCr - The things that most often break colors in practice are not looking at
MF_MT_YUV_MATRIXandMF_MT_VIDEO_NOMINAL_RANGE, and assuming the stride iswidth * bytesPerPixel
In short: if you want the easy path, have the Source Reader output RGB32. If you need high-volume processing or control over color, receive the frames as YUV and convert them yourself. Those are the two choices.
flowchart TB
accTitle: The two choices in this article
accDescr: Diagram showing the two choices this article covers - let the Source Reader output RGB32 if you want the easy path, and receive YUV and convert it yourself if you need high-volume processing or control over color.
want1["You want RGB frames"] -->|"Take the easy path"| pa1["Pattern A: let the Source Reader output RGB32"]
want1 -->|"High volume and color control"| pb1["Pattern B: receive YUV and convert it yourself"]
Figure 1: The choice is Pattern A if you take the easy path, and Pattern B if you take on the throughput and the responsibility for color.
In the diagram a solid line marks a relation that always holds and a dashed line marks a conditional one (the conditions are given per relation on the detail page). The full list of relations (23 in total, with evidence and certainty) and the definitions of the main concepts are collected on the knowledge map detail page (in Japanese). Data: JSON-LD / Turtle
2. Start with a Picture
It is faster to first look at a diagram of what happens inside Media Foundation.
flowchart LR
File["MP4 / H.264 / HEVC"] --> Decoder["decoder"]
Decoder --> YUV["YUV frames such as NV12 / YUY2 / YV12"]
YUV -->|Pattern A| SRVP["Source Reader video processing"]
SRVP --> RGB1["RGB32"]
YUV -->|Pattern B| App["Your own conversion code"]
App --> RGB2["BGRA / RGB"]
Figure 2: What the decoder emits is YUV frames, and from there the road to RGB forks into Pattern A and Pattern B.
If the contents of the video file are a compressed format such as H.264 or HEVC, the decoder first turns them back into uncompressed frames. These uncompressed frames are not necessarily RGB. In fact, in the Windows video world, YUV-family formats are the norm.
So when the application wants RGB, you choose one of the following.
- Have Media Foundation take the frames all the way to RGB32
- Receive YUV and turn it into RGB with your own code
This article is exactly about that fork in the road.
3. Sorting Out the Relationship Between YUV and RGB First
3.1. It Says YUV, but It Is Really About Y’CbCr
Windows API names and documentation broadly use the term YUV. In the context of digital video, however, you can read U as Cb and V as Cr with essentially no problem.
Roughly speaking:
Yis the brightness-oriented componentU/Vare the color-difference componentsRGBis each pixel directly carrying Red / Green / Blue
That is the relationship.
The human eye is more sensitive to fine detail in brightness than in color. So for video, a design that keeps Y at full detail and U/V somewhat coarser pays off. This is why YUV-family formats are so widely used.
flowchart TB
accTitle: Why YUV-family formats are used
accDescr: Diagram showing that the human eye is more sensitive to fine detail in brightness than in color, so a design that keeps Y fine and U/V coarse pays off, which is the reason YUV-family formats are widely used.
eye1["The human eye is sensitive to brightness"] --> dsn1["Keep Y fine and U/V coarse"]
dsn1 --> why1["Why YUV-family formats are used"]
Figure 3: The YUV-family design keeps Y fine and U/V coarse to match how sensitive the eye is to brightness.
3.2. 4:4:4 / 4:2:2 / 4:2:0 Is “How Much the Color Is Thinned Out”
This is the key to reading YUV.
| Notation | Meaning | Typical examples |
|---|---|---|
| 4:4:4 | Each pixel has its own Y/U/V | AYUV, I444 |
| 4:2:2 | 2 pixels horizontally share one U/V pair | YUY2, UYVY, I422 |
| 4:2:0 | A 2x2 pixel block shares one U/V pair | NV12, YV12, I420 |
It helps a lot to first look at the shape of the two formats you encounter most in practice.
Before that, let us pin down one term. The stride (also called the pitch) is the number of bytes in one row. It is not the image width itself: it is how many bytes you advance to reach the start of the next row, including any padding at the end of the row. This article uses stride and pitch with the same meaning. Microsoft Learn uses both wordings as well, so there is nothing to disambiguate.
In the diagrams below, W is the width, H is the height, and S is the stride. The point is that S >= W, and S == W is not guaranteed.
NV12 (4:2:0, planar) / width = W, height = H, stride = S
<----------- S bytes ----------->
<--- W --->
+-----------+---------------------+ --+
| Y Y Y Y Y | (padding) | |
| Y Y Y Y Y | (padding) | | Y plane
| Y Y Y Y Y | (padding) | | S * H bytes
| Y Y Y Y Y | (padding) | |
+-----------+---------------------+ --+ <- plane boundary = S * H from the start
| U V U V U | (padding) | |
| U V U V U | (padding) | | UV plane
+-----------+---------------------+ --+ height is H / 2 rows
Y of row y : yPlane + S * y
UV of row y : uvPlane + S * (y / 2)
Start of UV : scanline0 + S * H
In NV12, the 4 pixels of a 2x2 block share a single U/V pair. Y exists for each individual pixel.
The UV plane uses the same stride as the Y plane, but it has half as many rows. That is why the plane boundary is S * H and not W * H (we come back to this in 7.5.).
YUY2 (4:2:2, packed) / width = W, height = H, stride = S
<-------------- S bytes --------------->
<------- W * 2 bytes ------->
+-----------------------------+----------+
| Y0 U0 Y1 V0 Y2 U2 Y3 V2 … | (padding)| row 0
| Y0 U0 Y1 V0 Y2 U2 Y3 V2 … | (padding)| row 1
+-----------------------------+----------+
Start of row y : scanline0 + S * y
2 pixels = 4 bytes (Y, U, Y, V)
Only one plane (packed, so there is no boundary)
In YUY2, 2 horizontal pixels share one U/V pair. Y0 and Y1 are separate, but U0 and V0 are shared.
Being packed, there is no plane boundary to compute, but moving between rows still uses the stride.
At this point you can already see that YUV -> RGB is not a simple one-pixel-to-one-pixel substitution. First you have to think about how to assign the shared U/V to which pixels.
flowchart TB
accTitle: NV12 and YUY2 sharing units compared
accDescr: Diagram showing that NV12 shares one U/V pair across the 4 pixels of a 2x2 block while YUY2 shares one pair across 2 horizontal pixels, so you first have to work out how to assign the shared U/V to each pixel.
nv1["NV12 (4:2:0)"] --> sh1["4 pixels in a 2x2 block share one U/V pair"]
yy1["YUY2 (4:2:2)"] --> sh2["2 horizontal pixels share one U/V pair"]
sh1 --> asn1["Decide how to assign them to each pixel"]
sh2 --> asn1
Figure 4: Both formats share U/V, so the conversion is not a matter of substituting one pixel at a time.
3.3. YUV -> RGB Is “Color Space Conversion + Sampling Conversion”
If you look at Media Foundation’s Extended Color Information, strictly correct color conversion has quite a few stages: inverse quantization, chroma upsampling, YUV -> RGB, the transfer function, primaries conversion, and finally quantization.
That said, as practical code for 8-bit SDR, it is easiest to understand if you split it into the following three layers.
- Undo the subsampling Expand the 4:2:0 or 4:2:2 U/V so that every pixel can reference a value
- Undo the range Video Y normally uses 16..235 and U/V use 16..240, so undo that scaling
- Apply the matrix
Convert to RGB using coefficients such as
BT.601orBT.709
In other words, in practical terms, YUV -> RGB conversion is the process of deciding:
- which U/V is the color for that pixel
- which coefficients to use to turn that Y/U/V back into RGB
flowchart TB
accTitle: The three layers to get right in practical code
accDescr: Diagram showing that in practical 8-bit SDR code it is easier to understand YUV to RGB conversion as three layers - undo the subsampling, undo the range, and apply the matrix.
y1["YUV frame"] --> up1["Undo the subsampling"]
up1 --> rg1["Undo the range (16..235 and so on)"]
rg1 --> mx1["Apply the matrix (601 / 709)"]
mx1 --> rgb2["RGB"]
Figure 5: The conversion is not three coefficients; it is built from three layers - subsampling, range, and matrix.
3.4. Treat BT.601 and BT.709 Carelessly and Colors Drift Subtly
The Media Foundation documentation describes the relationship as BT.601 being preferred for SDTV and below, and BT.709 for video beyond SD.
However, silently guessing “the resolution is large, so it must be 709” is not a great idea. Color drift does not crash, so it easily slips into production unnoticed.
Media Foundation can carry color space information as media type attributes. At minimum, look at these two:
MF_MT_YUV_MATRIXMF_MT_VIDEO_NOMINAL_RANGE
Looking at these two and explicitly accepting only the combinations your code supports makes it far less likely that something breaks silently later.
flowchart TB
accTitle: A flow that avoids guessing the color space
accDescr: Diagram showing that silently guessing 601 or 709 from the resolution lets color drift slip into production unnoticed, so the flow is to look at MF_MT_YUV_MATRIX and MF_MT_VIDEO_NOMINAL_RANGE and explicitly accept only the combinations you support.
gs1["Silently guess from the resolution"] --> sl1["Color drift does not crash"]
sl1 --> op1["It slips into production unnoticed"]
at1["Look at the matrix and range attributes"] --> ps1["Accept only the combinations you support"]
ps1 -.-> sf1["Silent breakage is prevented"]
Figure 6: Do not guess the color space - read the attributes and explicitly accept only the combinations you can handle.
3.5. The First Formula to Learn Is the BT.601 Limited-Range Version
The canonical 8-bit BT.601 formula looks like this.
C = Y - 16
D = U - 128
E = V - 128
R = clip(1.164383 * C + 1.596027 * E)
G = clip(1.164383 * C - 0.391762 * D - 0.812968 * E)
B = clip(1.164383 * C + 2.017232 * D)
For BT.709 the coefficients change. We will show that in code later.
What matters here is not memorizing the coefficients but the structure: subtract the black level 16 from Y, and view U/V as centered on 128.
flowchart TB
accTitle: The structure of the conversion formula
accDescr: Diagram showing the structure of the conversion formula - subtract the black level of 16 from Y, view U and V as centered on 128, multiply by the coefficients for each matrix, and clip the result.
yy2["Subtract 16 from Y (black level)"] --> co1["Multiply by the matrix coefficients"]
uv1["Subtract 128 from U / V (center)"] --> co1
co1 --> cl1["Clip to 0..255"]
Figure 7: What you memorize is not the coefficients but the structure of the formula - black level 16 and center 128.
4. Pattern A: Let Media Foundation Convert Automatically
4.1. When This Is a Good Fit
This approach is well suited to situations like the following.
- You want to extract a single still image from an MP4
- You want to create a few thumbnails
- You want an RGB image to hand to WIC
- Batch or tooling use is fine; this is not real-time playback
The Source Reader has a feature that performs limited video processing of YUV -> RGB32 when you use MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING.
However, as Microsoft Learn also notes, this is software processing and not optimized for playback. If you want to process hundreds of frames per second, leaning on this is not quite the right tool.
flowchart TB
accTitle: What the automatic conversion suits and does not suit
accDescr: Diagram showing that the Source Reader automatic conversion is software processing that is not optimized for playback, so it suits still-image extraction, thumbnails, and batch work, while leaning on it for hundreds of frames per second is the wrong tool.
auto1["Source Reader automatic conversion"] --> sw1["Software processing"]
sw1 -->|"Good fit"| bat1["Stills, thumbnails, batch work"]
sw1 -.->|"Do not lean on it"| rt1["Hundreds of frames per second in real time"]
Figure 8: The automatic conversion is software processing, so keep it to tooling that handles a small number of frames.
4.2. What You Set to Get RGB32 Out
The flow is quite straightforward.
- In the attributes passed to
MFCreateSourceReaderFromURL, setMF_SOURCE_READER_ENABLE_VIDEO_PROCESSING = TRUE - Select the video stream
- Request
MFMediaType_Video/MFVideoFormat_RGB32viaSetCurrentMediaType - Read samples with
ReadSample
That alone makes the limited video processing inserted behind the decoder do the YUV -> RGB32 for you.
flowchart TB
accTitle: The four steps that turn on automatic conversion
accDescr: Diagram showing the four steps of automatic conversion - enable video processing in the attributes and create the reader, select the video stream, request RGB32, and read with ReadSample.
a1["Enable video processing in the attributes"] --> a2["Select the video stream"]
a2 --> a3["Request RGB32"]
a3 --> a4["Read with ReadSample"]
a4 -.-> a5["Converted to RGB32 behind the decoder"]
Figure 9: Follow those four steps and the video processing behind the decoder carries the frames all the way to RGB32.
4.3. Code
The following code assumes CoInitializeEx and MFStartup have already been done. A minimal version looks roughly like this.
#include <windows.h>
#include <mfapi.h>
#include <mfidl.h>
#include <mfreadwrite.h>
#include <mferror.h>
#include <wrl/client.h>
#pragma comment(lib, "mfplat.lib")
#pragma comment(lib, "mfreadwrite.lib")
#pragma comment(lib, "mfuuid.lib")
#pragma comment(lib, "ole32.lib")
using Microsoft::WRL::ComPtr;
HRESULT CreateSourceReaderWithAutoRgb(
const wchar_t* path,
IMFSourceReader** ppReader)
{
if (!path || !ppReader) return E_POINTER;
*ppReader = nullptr;
ComPtr<IMFAttributes> attrs;
HRESULT hr = MFCreateAttributes(&attrs, 2);
if (FAILED(hr)) return hr;
hr = attrs->SetUINT32(MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING, TRUE);
if (FAILED(hr)) return hr;
hr = MFCreateSourceReaderFromURL(path, attrs.Get(), ppReader);
if (FAILED(hr)) return hr;
hr = (*ppReader)->SetStreamSelection(MF_SOURCE_READER_ALL_STREAMS, FALSE);
if (FAILED(hr)) return hr;
hr = (*ppReader)->SetStreamSelection(MF_SOURCE_READER_FIRST_VIDEO_STREAM, TRUE);
if (FAILED(hr)) return hr;
ComPtr<IMFMediaType> outType;
hr = MFCreateMediaType(&outType);
if (FAILED(hr)) return hr;
hr = outType->SetGUID(MF_MT_MAJOR_TYPE, MFMediaType_Video);
if (FAILED(hr)) return hr;
hr = outType->SetGUID(MF_MT_SUBTYPE, MFVideoFormat_RGB32);
if (FAILED(hr)) return hr;
hr = (*ppReader)->SetCurrentMediaType(
MF_SOURCE_READER_FIRST_VIDEO_STREAM,
nullptr,
outType.Get());
if (FAILED(hr)) return hr;
return S_OK;
}
HRESULT ReadOneRgb32Sample(
IMFSourceReader* reader,
IMFSample** ppSample,
LONGLONG* pTimestamp100ns)
{
if (!reader || !ppSample) return E_POINTER;
*ppSample = nullptr;
if (pTimestamp100ns) *pTimestamp100ns = 0;
DWORD streamIndex = 0;
DWORD flags = 0;
LONGLONG timestamp = 0;
HRESULT hr = reader->ReadSample(
MF_SOURCE_READER_FIRST_VIDEO_STREAM,
0,
&streamIndex,
&flags,
×tamp,
ppSample);
if (FAILED(hr)) return hr;
if (flags & MF_SOURCE_READERF_ENDOFSTREAM) return MF_E_END_OF_STREAM;
if (*ppSample == nullptr) return MF_E_INVALID_STREAM_DATA;
if (pTimestamp100ns) *pTimestamp100ns = timestamp;
return S_OK;
}
After this, calling GetCurrentMediaType lets you check the actual output size and stride.
4.4. Strengths of This Approach
The good thing about this approach is that it gets you to a correct picture quickly.
- You do not have to write the 4:2:0 / 4:2:2 expansion yourself
- It hides much of the hassle of matrix handling / deinterlacing
- The output is easy to hand to WIC or GDI
- For processing a handful of frames, it is perfectly practical
For still-image extraction tools, starting here is quite natural.
4.5. But There Are Pitfalls Too
This automatic conversion has the following characteristics.
| Item | Details |
|---|---|
| Conversion target | Basically RGB32 |
| Implementation | Software processing |
| Suited for | Small numbers of frames, thumbnails, offline processing |
| Not suited for | D3D-based real-time rendering, high-volume frame processing |
| Incompatible attributes | MF_SOURCE_READER_D3D_MANAGER, MF_READWRITE_DISABLE_CONVERTERS |
And one more important thing: the handling of the 4th byte in RGB32.
In memory, Windows RGB32 is laid out as Blue / Green / Red / Alpha or Don’t Care. It is not ARGB32. If you pass it to WIC as 32bppBGRA, it is safer to fill the 4th byte with 0xFF to make it opaque.
flowchart TB
accTitle: Handling the 4th byte of RGB32
accDescr: Diagram showing that the 4th byte after B, G, and R in Windows RGB32 is either alpha or don't care, so filling it with 0xFF to make it opaque before handing the data to WIC as 32bppBGRA is safer.
r32["RGB32 memory layout"] --> bgr1["3 bytes of B, G, R"]
r32 --> b41["The 4th byte is alpha or don't care"]
b41 -->|"Fill it with 0xFF"| wic1["Can be handed to WIC as 32bppBGRA"]
b41 -.->|"Hand it over as is"| tr3["It can come out transparent"]
Figure 10: The 4th byte is undefined, so make it opaque with 0xFF before handing the data to WIC.
We touched on this as an easy thing to trip over in the previous still-image extraction article as well.
5. Pattern B: Write the Conversion Yourself
5.1. When This Is a Good Fit
Doing the conversion yourself is a good fit in cases like these.
- You process a large number of frames and want to optimize the conversion yourself
- You want to feed
NV12straight to the GPU or SIMD code - You want to handle
BT.601/BT.709/ range explicitly - You want to produce output formats other than
RGB32 - The Source Reader’s limited automatic conversion is not enough
You could call it the pattern where you take on responsibility for throughput and color in exchange for freedom.
flowchart TB
accTitle: The trade-off of manual conversion
accDescr: Diagram showing that manual conversion takes on responsibility for throughput and color in exchange for optimization, a path into GPU and SIMD code, explicit control of matrix and range, and freedom in the output format.
own1["Choose manual conversion"] --> res1["You own throughput and color"]
res1 --> fr1["Freedom in optimization, GPU / SIMD, and output format"]
res1 --> fr2["Explicit control of matrix / range"]
Figure 11: Manual conversion trades responsibility for performance and freedom over color.
5.2. Overall Flow of Manual Conversion
The steps are as follows.
- Set the Source Reader output to
NV12orYUY2 - Get the actual subtype and attributes with
GetCurrentMediaType - Check
MF_MT_FRAME_SIZE,MF_MT_DEFAULT_STRIDE,MF_MT_YUV_MATRIX, andMF_MT_VIDEO_NOMINAL_RANGE - Take the buffer out of the sample and lock it
- Work out the Y/U/V that each pixel references
- Apply the matrix and write BGRA
The code in this article narrows the scope to 8-bit SDR / progressive / NV12 or YUY2 / limited range.
Narrowing the assumptions here is not laziness; it actually matters. A YUV converter written to “accept everything for now” tends to break colors silently.
flowchart TB
accTitle: Overall flow of manual conversion
accDescr: Diagram showing the manual conversion steps - request NV12 or YUY2, check the actual media type and attributes, lock the buffer, work out the Y/U/V for each pixel, and apply the matrix to write BGRA.
f1["Request NV12 / YUY2"] --> f2["Check the actual subtype and attributes"]
f2 --> f3["Lock the buffer"]
f3 --> f4["Work out the Y/U/V for each pixel"]
f4 --> f5["Apply the matrix and write BGRA"]
f2 -.-> nar1["Narrowing the assumptions keeps colors from breaking"]
Figure 12: Manual conversion runs in the order request, check, lock, reference, convert, and the narrower the assumptions the safer it gets.
5.3. First, Specify the Output Media Type Explicitly
First, tell the Source Reader that you want the YUV as-is. This also assumes CoInitializeEx / MFStartup have already been done.
#include <windows.h>
#include <mfapi.h>
#include <mfidl.h>
#include <mfreadwrite.h>
#include <mferror.h>
#include <wrl/client.h>
using Microsoft::WRL::ComPtr;
HRESULT ConfigureSourceReaderForSubtype(
IMFSourceReader* reader,
REFGUID subtype)
{
if (!reader) return E_POINTER;
HRESULT hr = reader->SetStreamSelection(MF_SOURCE_READER_ALL_STREAMS, FALSE);
if (FAILED(hr)) return hr;
hr = reader->SetStreamSelection(MF_SOURCE_READER_FIRST_VIDEO_STREAM, TRUE);
if (FAILED(hr)) return hr;
ComPtr<IMFMediaType> outType;
hr = MFCreateMediaType(&outType);
if (FAILED(hr)) return hr;
hr = outType->SetGUID(MF_MT_MAJOR_TYPE, MFMediaType_Video);
if (FAILED(hr)) return hr;
hr = outType->SetGUID(MF_MT_SUBTYPE, subtype);
if (FAILED(hr)) return hr;
hr = reader->SetCurrentMediaType(
MF_SOURCE_READER_FIRST_VIDEO_STREAM,
nullptr,
outType.Get());
if (FAILED(hr)) return hr;
return S_OK;
}
Here you pass either MFVideoFormat_NV12 or MFVideoFormat_YUY2 as subtype.
What you need to watch out for is that the subtype you requested does not necessarily go through. Confirm what actually comes out with GetCurrentMediaType.
flowchart TB
accTitle: Separating what you request from what actually comes out
accDescr: Diagram showing that the subtype requested through SetCurrentMediaType does not necessarily go through, so the flow is to confirm what actually comes out with GetCurrentMediaType.
req1["Request a subtype"] -.->|"May not go through as is"| out2["The actual output"]
out2 --> gct1["Confirm it with GetCurrentMediaType"]
gct1 --> use1["Write the rest of the code against the confirmed values"]
Figure 13: A request is only a request - always confirm the actual output with GetCurrentMediaType before you use it.
5.4. Before Converting, Accept Only the Color Information You Support
For a manual conversion, first pull the minimum information from the media type.
The sample in this article accepts only NV12 / YUY2, and lets through only BT.601 or BT.709 for the matrix and only MFNominalRange_16_235 for the range.
#include <vector>
struct DecodedFrameInfo
{
GUID subtype = GUID_NULL;
UINT32 width = 0;
UINT32 height = 0;
LONG defaultStride = 0;
MFVideoTransferMatrix matrix = MFVideoTransferMatrix_Unknown;
MFNominalRange nominalRange = MFNominalRange_Unknown;
};
HRESULT GetDefaultStride(
IMFMediaType* pType,
LONG* plStride)
{
if (!pType || !plStride) return E_POINTER;
LONG stride = 0;
HRESULT hr = pType->GetUINT32(
MF_MT_DEFAULT_STRIDE,
reinterpret_cast<UINT32*>(&stride));
if (FAILED(hr))
{
GUID subtype = GUID_NULL;
UINT32 width = 0;
UINT32 height = 0;
hr = pType->GetGUID(MF_MT_SUBTYPE, &subtype);
if (FAILED(hr)) return hr;
hr = MFGetAttributeSize(pType, MF_MT_FRAME_SIZE, &width, &height);
if (FAILED(hr)) return hr;
hr = MFGetStrideForBitmapInfoHeader(subtype.Data1, width, &stride);
if (FAILED(hr)) return hr;
hr = pType->SetUINT32(MF_MT_DEFAULT_STRIDE, static_cast<UINT32>(stride));
if (FAILED(hr)) return hr;
}
*plStride = stride;
return S_OK;
}
HRESULT GetStrictDecodedFrameInfo(
IMFMediaType* pType,
DecodedFrameInfo* pInfo)
{
if (!pType || !pInfo) return E_POINTER;
HRESULT hr = pType->GetGUID(MF_MT_SUBTYPE, &pInfo->subtype);
if (FAILED(hr)) return hr;
if (pInfo->subtype != MFVideoFormat_NV12 &&
pInfo->subtype != MFVideoFormat_YUY2)
{
return MF_E_INVALIDMEDIATYPE;
}
hr = MFGetAttributeSize(pType, MF_MT_FRAME_SIZE, &pInfo->width, &pInfo->height);
if (FAILED(hr)) return hr;
hr = GetDefaultStride(pType, &pInfo->defaultStride);
if (FAILED(hr)) return hr;
UINT32 value = 0;
hr = pType->GetUINT32(MF_MT_YUV_MATRIX, &value);
if (FAILED(hr)) return hr;
pInfo->matrix = static_cast<MFVideoTransferMatrix>(value);
if (pInfo->matrix != MFVideoTransferMatrix_BT601 &&
pInfo->matrix != MFVideoTransferMatrix_BT709)
{
return MF_E_INVALIDMEDIATYPE;
}
hr = pType->GetUINT32(MF_MT_VIDEO_NOMINAL_RANGE, &value);
if (FAILED(hr)) return hr;
pInfo->nominalRange = static_cast<MFNominalRange>(value);
if (pInfo->nominalRange != MFNominalRange_16_235)
{
return MF_E_INVALIDMEDIATYPE;
}
return S_OK;
}
This is deliberately strict.
The Media Foundation enum documentation does say things like “treat Unknown as BT.709,” but in practice, silently rounding here makes color drift hard to notice. At least in a first implementation, returning an error for combinations you do not support is safer.
When Does Unknown Come Back?
In case this feels too strict, here are the paths that produce Unknown. They are mostly cases where the source video simply does not carry color information.
- The H.264 / HEVC VUI has no color description. Per the specification, when
colour_description_present_flagis 0,matrix_coefficientsis treated as unspecified. If that missing information goes through the decoder, the matrix passed downstream is unspecified as well - Raw YUV from a capture device or an old container. These are paths that carry no color space description at all
- Sometimes the
MF_MT_YUV_MATRIXattribute is not present at all. In that caseGetUINT32returns no value and fails withMF_E_ATTRIBUTENOTFOUND(the code above rejects it right there viaFAILED(hr))
The important point is that Unknown does not mean “we know it is BT.709”; it means “we do not know.” Applying 709 to SD material drifts the colors, and so does the reverse.
From there, the approach splits in two.
- Reject it strictly (the approach in this article): return an error as unsupported and let the layer above decide that this material is not handled. Saying outright that you cannot handle it is safer than letting colors drift quietly
- Pick a default and let it through: if you really have to let it through, write down in the log what you assumed when the value was
Unknown. Then state explicitly that you picked 601 / 709 based on the resolution
Either way, the one thing to avoid is silently rounding it off. Color drift does not crash, so it rides into production unnoticed.
flowchart TB
accTitle: What to do when the matrix is Unknown
accDescr: Diagram showing that Unknown means you do not know rather than knowing it is BT.709, so the approach splits into rejecting it strictly with an error or letting it through with the assumption written to the log, and that silently rounding it off is the one thing to avoid.
unk1["matrix is Unknown = we do not know"] -->|"The approach in this article"| st4["Reject it strictly with an error"]
unk1 -->|"If you have to let it through"| dflt1["Log what you assumed and let it through"]
unk1 -.->|"The one thing to avoid"| mute1["Round it off silently"]
Figure 14: Unknown means unknown, so either reject it or let it through with a log entry - never round it off silently.
With cameras and JPEG-family sources, you sometimes want to handle full-range paths separately. Rather than quietly serving both here, the approach is to explicitly narrow the assumptions this code accepts.
5.5. Read the Buffer Trusting the Stride
This part is quite important too.
MF_MT_DEFAULT_STRIDEis the minimum stride- The actual sample buffer may have an actual stride that includes padding
- If
IMF2DBuffer::Lock2Dis available, prefer it
Taking the helper pattern from Microsoft Learn’s Uncompressed Video Buffers and making it directly usable gives us this.
class BufferLock
{
public:
explicit BufferLock(IMFMediaBuffer* buffer)
: m_buffer(buffer),
m_2dBuffer(nullptr),
m_locked(false)
{
if (m_buffer)
{
m_buffer->AddRef();
m_buffer->QueryInterface(IID_PPV_ARGS(&m_2dBuffer));
}
}
~BufferLock()
{
Unlock();
if (m_2dBuffer)
{
m_2dBuffer->Release();
m_2dBuffer = nullptr;
}
if (m_buffer)
{
m_buffer->Release();
m_buffer = nullptr;
}
}
HRESULT Lock(
LONG defaultStride,
DWORD heightInPixels,
BYTE** ppScanline0,
LONG* pActualStride)
{
if (!m_buffer || !ppScanline0 || !pActualStride) return E_POINTER;
if (m_locked) return MF_E_INVALIDREQUEST;
if (m_2dBuffer)
{
HRESULT hr = m_2dBuffer->Lock2D(ppScanline0, pActualStride);
if (FAILED(hr)) return hr;
m_locked = true;
return S_OK;
}
BYTE* pData = nullptr;
HRESULT hr = m_buffer->Lock(&pData, nullptr, nullptr);
if (FAILED(hr)) return hr;
*pActualStride = defaultStride;
if (defaultStride < 0)
{
*ppScanline0 =
pData + static_cast<size_t>(-defaultStride) * (heightInPixels - 1);
}
else
{
*ppScanline0 = pData;
}
m_locked = true;
return S_OK;
}
void Unlock()
{
if (!m_locked) return;
if (m_2dBuffer)
{
m_2dBuffer->Unlock2D();
}
else
{
m_buffer->Unlock();
}
m_locked = false;
}
private:
IMFMediaBuffer* m_buffer;
IMF2DBuffer* m_2dBuffer;
bool m_locked;
};
The recommended YUV surface definitions use top-left origin and a positive stride, but for actual buffer access it is safer to use the stride (that is, the pitch) the API handed back. Hard-code a width-based value here and things break silently later.
flowchart TB
accTitle: Stride priority order
accDescr: Diagram showing the priority order for stride - MF_MT_DEFAULT_STRIDE is the minimum stride, the actual buffer may carry padding, so prefer the value Lock2D returns when it is available and avoid hard-coding a width-based value.
p1["Actual stride returned by Lock2D"] -->|"First choice when available"| acc1["The value used for buffer access"]
p2["MF_MT_DEFAULT_STRIDE (the minimum)"] -->|"fallback"| acc1
p3["Hard-coded width-based value"] -.->|"Breaks silently"| acc1
Figure 15: For moving between rows, prefer the stride Lock2D actually reports and never derive it from the width.
5.6. Turning the Per-Pixel Conversion Formula into Code
Here we handle only the limited range of BT.601 and BT.709. The output is BGRA32, which is easy to hand to WIC or GDI.
inline BYTE ClampToByte(double value)
{
if (value <= 0.0) return 0;
if (value >= 255.0) return 255;
return static_cast<BYTE>(value + 0.5);
}
HRESULT ConvertLimitedYuvPixelToBgra(
BYTE y,
BYTE u,
BYTE v,
MFVideoTransferMatrix matrix,
BYTE* dstPixel)
{
if (!dstPixel) return E_POINTER;
const double c = static_cast<double>(y) - 16.0;
const double d = static_cast<double>(u) - 128.0;
const double e = static_cast<double>(v) - 128.0;
double r = 0.0;
double g = 0.0;
double b = 0.0;
switch (matrix)
{
case MFVideoTransferMatrix_BT601:
r = 1.164383 * c + 1.596027 * e;
g = 1.164383 * c - 0.391762 * d - 0.812968 * e;
b = 1.164383 * c + 2.017232 * d;
break;
case MFVideoTransferMatrix_BT709:
r = 1.164383 * c + 1.792741 * e;
g = 1.164383 * c - 0.213249 * d - 0.532909 * e;
b = 1.164383 * c + 2.112402 * d;
break;
default:
return MF_E_INVALIDMEDIATYPE;
}
dstPixel[0] = ClampToByte(b);
dstPixel[1] = ClampToByte(g);
dstPixel[2] = ClampToByte(r);
dstPixel[3] = 255;
return S_OK;
}
What this does is simple.
- Subtract 16 from
Y - Subtract 128 from
U/V - Multiply by the coefficients for the given matrix
- Clip the result to 0..255
- Set the 4th BGRA byte to
255
5.7. Converting NV12 to BGRA32
NV12 is 4:2:0, so the 4 pixels of a 2x2 block share the same U/V.
As a minimal implementation, the most understandable approach is to use that shared chroma directly for all 4 pixels.
HRESULT ConvertNv12ToBgra32(
IMFMediaBuffer* buffer,
const DecodedFrameInfo& info,
std::vector<BYTE>& dstBgra)
{
if (!buffer) return E_POINTER;
if (info.subtype != MFVideoFormat_NV12) return MF_E_INVALIDMEDIATYPE;
if ((info.width & 1u) != 0 || (info.height & 1u) != 0)
{
return MF_E_INVALIDMEDIATYPE;
}
dstBgra.resize(static_cast<size_t>(info.width) * info.height * 4);
BufferLock lock(buffer);
BYTE* scanline0 = nullptr;
LONG actualStride = 0;
HRESULT hr = lock.Lock(
info.defaultStride,
info.height,
&scanline0,
&actualStride);
if (FAILED(hr)) return hr;
if (actualStride <= 0)
{
lock.Unlock();
return MF_E_INVALIDMEDIATYPE;
}
const BYTE* yPlane = scanline0;
// The UV plane starts stride * height bytes into the buffer.
// Note that it is not width * height (see the diagram in 3.2.)
const BYTE* uvPlane =
scanline0 + static_cast<size_t>(actualStride) * info.height;
for (UINT32 y = 0; y < info.height; ++y)
{
// Always move between rows in units of the stride
const BYTE* yRow = yPlane + static_cast<size_t>(actualStride) * y;
// 4:2:0, so 2 vertical rows share one UV row -> y / 2
// The UV plane uses the same stride as the Y plane
const BYTE* uvRow = uvPlane + static_cast<size_t>(actualStride) * (y / 2);
// The output is tightly packed BGRA with no padding, so width * 4
BYTE* dstRow =
dstBgra.data() + static_cast<size_t>(info.width) * 4 * y;
for (UINT32 x = 0; x < info.width; ++x)
{
const BYTE Y = yRow[x];
// The UV plane alternates [U, V].
// 2 horizontal pixels share one pair, so (x / 2) gives the pair index,
// and since one pair is 2 bytes, * 2 turns it into a byte offset. +0 is U, +1 is V.
// x = 0, 1 -> uvRow[0], uvRow[1]
// x = 2, 3 -> uvRow[2], uvRow[3]
const BYTE U = uvRow[(x / 2) * 2 + 0];
const BYTE V = uvRow[(x / 2) * 2 + 1];
hr = ConvertLimitedYuvPixelToBgra(
Y,
U,
V,
info.matrix,
dstRow + static_cast<size_t>(x) * 4);
if (FAILED(hr))
{
lock.Unlock();
return hr;
}
}
}
lock.Unlock();
return S_OK;
}
This code interprets the chroma upsampling in a nearest-neighbor fashion. Visually that is often perfectly serviceable, but if you are aiming for maximum quality, a design that first performs the 4:2:0 -> 4:2:2 -> 4:4:4 upconversion, as described in the Microsoft Learn YUV article, is theoretically cleaner.
flowchart TB
accTitle: Two designs for chroma upsampling
accDescr: Diagram showing that the minimal implementation reuses the shared chroma for all 4 pixels in a nearest-neighbor fashion and is often perfectly serviceable, while upconverting from 4:2:0 through 4:2:2 to 4:4:4 before converting is theoretically cleaner when image quality comes first.
min1["Minimal: use the shared chroma as is"] --> pr1["Often perfectly serviceable visually"]
hq1["Upconvert first, then convert"] --> pr2["Theoretically cleaner, quality first"]
hq1 -.-> steps1["In the order 4:2:0 → 4:2:2 → 4:4:4"]
Figure 16: Choosing between the minimal implementation that reuses the shared chroma and a design that upconverts in stages.
5.8. Converting YUY2 to BGRA32
YUY2 is packed 4:2:2.
Two pixels simply share one U/V pair, so it is a bit easier to read than NV12.
#include <cstddef>
HRESULT ConvertYuy2ToBgra32(
IMFMediaBuffer* buffer,
const DecodedFrameInfo& info,
std::vector<BYTE>& dstBgra)
{
if (!buffer) return E_POINTER;
if (info.subtype != MFVideoFormat_YUY2) return MF_E_INVALIDMEDIATYPE;
if ((info.width & 1u) != 0) return MF_E_INVALIDMEDIATYPE;
dstBgra.resize(static_cast<size_t>(info.width) * info.height * 4);
BufferLock lock(buffer);
BYTE* scanline0 = nullptr;
LONG actualStride = 0;
HRESULT hr = lock.Lock(
info.defaultStride,
info.height,
&scanline0,
&actualStride);
if (FAILED(hr)) return hr;
for (UINT32 y = 0; y < info.height; ++y)
{
const BYTE* src =
scanline0 +
static_cast<ptrdiff_t>(actualStride) * static_cast<ptrdiff_t>(y);
BYTE* dstRow =
dstBgra.data() + static_cast<size_t>(info.width) * 4 * y;
for (UINT32 x = 0; x < info.width; x += 2)
{
const BYTE Y0 = src[0];
const BYTE U = src[1];
const BYTE Y1 = src[2];
const BYTE V = src[3];
hr = ConvertLimitedYuvPixelToBgra(
Y0,
U,
V,
info.matrix,
dstRow + static_cast<size_t>(x) * 4);
if (FAILED(hr))
{
lock.Unlock();
return hr;
}
hr = ConvertLimitedYuvPixelToBgra(
Y1,
U,
V,
info.matrix,
dstRow + static_cast<size_t>(x + 1) * 4);
if (FAILED(hr))
{
lock.Unlock();
return hr;
}
src += 4;
}
}
lock.Unlock();
return S_OK;
}
Because YUY2 lays the bytes out as Y0 U Y1 V, the structure of “reuse the U/V for every 2 pixels” is visible directly in the data.
That makes the mental model easier to build than for NV12.
5.9. The Entry Point When Calling from a Sample
Finally, pulling a contiguous buffer out of the IMFSample and branching on the subtype makes this easy to use.
HRESULT ConvertSampleToBgra32(
IMFSample* sample,
const DecodedFrameInfo& info,
std::vector<BYTE>& dstBgra)
{
if (!sample) return E_POINTER;
ComPtr<IMFMediaBuffer> buffer;
HRESULT hr = sample->ConvertToContiguousBuffer(&buffer);
if (FAILED(hr)) return hr;
if (info.subtype == MFVideoFormat_NV12)
{
return ConvertNv12ToBgra32(buffer.Get(), info, dstBgra);
}
if (info.subtype == MFVideoFormat_YUY2)
{
return ConvertYuy2ToBgra32(buffer.Get(), info, dstBgra);
}
return MF_E_INVALIDMEDIATYPE;
}
With that, the earlier stage becomes
- create the reader
- request
NV12orYUY2 - build a
DecodedFrameInfofromGetCurrentMediaType ReadSampleConvertSampleToBgra32
as a flow.
flowchart TB
accTitle: Branching in the entry-point function
accDescr: Diagram showing the structure of the entry-point function - pull a contiguous buffer out of the sample, branch to the NV12 conversion when the subtype is NV12 and to the YUY2 conversion when it is YUY2, and return an error for anything else.
smp1["IMFSample"] --> cont1["Pull out a contiguous buffer"]
cont1 -->|"NV12"| cnv1["To the NV12 conversion"]
cont1 -->|"YUY2"| cnv2["To the YUY2 conversion"]
cont1 -->|"Anything else"| err1["Return an error"]
Figure 17: Make the buffer contiguous at the entry point, branch per subtype, and let anything unsupported fall straight through to an error.
The actual calling side looks something like this.
ComPtr<IMFMediaType> currentType;
HRESULT hr = reader->GetCurrentMediaType(
MF_SOURCE_READER_FIRST_VIDEO_STREAM,
¤tType);
if (FAILED(hr)) return hr;
DecodedFrameInfo info;
hr = GetStrictDecodedFrameInfo(currentType.Get(), &info);
if (FAILED(hr)) return hr;
DWORD flags = 0;
LONGLONG timestamp = 0;
ComPtr<IMFSample> sample;
hr = reader->ReadSample(
MF_SOURCE_READER_FIRST_VIDEO_STREAM,
0,
nullptr,
&flags,
×tamp,
&sample);
if (FAILED(hr)) return hr;
if (flags & MF_SOURCE_READERF_ENDOFSTREAM) return MF_E_END_OF_STREAM;
if (!sample) return MF_E_INVALID_STREAM_DATA;
std::vector<BYTE> bgra;
hr = ConvertSampleToBgra32(sample.Get(), info, bgra);
if (FAILED(hr)) return hr;
// bgra can be treated as top-down / 32bpp BGRA
5.10. Where to Put the “Manual Conversion”
The code so far takes the form of the application converting after the Source Reader. That is the easiest to understand.
However, if you want to insert the conversion inside the Media Foundation pipeline, there are other designs.
- Write your own
MFT - Use the
Video Processor MFT/ XVP - Write an
NV12-> RGB shader on the GPU side
Going that far changes the topic somewhat, so this article focused on application-side code. Still, it is useful to know that between “let Media Foundation handle it” and “do everything in the app,” there is a middle ground: the Video Processor MFT.
5.11. Confirming That the Conversion Is Correct
Color bugs are hard to see, so check “it runs” and “it is correct” separately. There are two stages, in this order.
Stage 1: Feed in Known Values and Compare Against Hand Calculation
Rather than starting with a video, it is more reliable to pass known Y/U/V values to ConvertLimitedYuvPixelToBgra. You need neither a video file nor Media Foundation.
For BT.601 limited range, here are the Y/U/V values for representative colors and the values you should expect from the formula in 5.6.
| Color | Y | U | V | Expected R | G | B |
|---|---|---|---|---|---|---|
| Black | 16 | 128 | 128 | 0 | 0 | 0 |
| White | 235 | 128 | 128 | 255 | 255 | 255 |
| Red | 81 | 90 | 240 | 254 | 0 | 0 |
| Blue | 41 | 240 | 110 | 0 | 0 | 255 |
For red, for example, put C = 81 - 16 = 65, D = 90 - 128 = -38, and E = 240 - 128 = 112 into the formula, and you get
R = 1.164383 * 65 + 1.596027 * 112 = 254.44 -> 254
G = 1.164383 * 65 - 0.391762 * (-38)
- 0.812968 * 112 = -0.48 -> 0
B = 1.164383 * 65 + 2.017232 * (-38) = -0.97 -> 0
The output is in BGRA order, so as a byte sequence it is 00 00 FE FF.
What matters here is that red comes out as 254 rather than 255. The reason is not the precision of the coefficients. It is that the input Y/U/V are already rounded integers.
Taking the theoretical red (255, 0, 0) down into BT.601 limited range gives Y = 16 + 219 × 0.299 = 81.481, U = 90.203, and V lands exactly on 240. The moment you store that as an 8-bit sample, the fraction disappears and Y becomes 81. The lost 0.481 turns into a shortfall of 0.481 × 1.164383 ≈ 0.56 on the way back. 255 - 0.56 = 254.44 - that is where the 254.44 above comes from. Even with infinite-precision coefficients the result stays 254.44; rounding to 6 decimal places only shows up below the fourth decimal place and never reaches the 8-bit output.
How you finally convert to integers also affects the result. ClampToByte in 5.6. constrains the value to [0, 255] and then truncates value + 0.5, which is round-half-up. With plain truncation (static_cast<BYTE>(value)), this red is still 254, but values sitting near a boundary, such as blue’s B = 255.04 or red’s R = 0.38, come out 1 off. Before you compare against another implementation, find out which one it uses.
In other words, the reason to allow a difference of 1 or 2 is not “the coefficients have different precision” but the two facts that “sampling dropped a fraction” and “the rounding policy differs between implementations.” Put the other way around, a difference those two cannot explain is a real bug. Red coming out as 250, red and blue swapped, shadows lifted - for differences like these, suspect the assumptions of the conversion (BT.601 mistaken for BT.709, full range mistaken for limited range, U and V swapped, a misread stride) rather than the precision of the coefficients. Write these off as “a precision issue” and you miss bugs you could have fixed.
flowchart TB
accTitle: Drawing the line between tolerance and a real bug
accDescr: Diagram showing that a difference of one or two can be explained by the fraction dropped in sampling and by the differing rounding policies, and that a difference those two cannot explain is a real bug whose conversion assumptions should be suspected.
dif1["Look at the difference from the expected value"] -->|"A difference of 1 to 2"| exp1["Explained by the sampling fraction and the rounding policy"]
dif1 -->|"A difference with no explanation"| bug1["A real bug"]
bug1 --> prem1["Suspect the matrix, range, U/V, and stride assumptions"]
Figure 18: Small differences are explained by sampling and rounding; a difference they cannot explain points to a mistaken assumption.
Written as a test, this shape is enough.
#include <cstdlib> // std::abs
// Check whether the difference from the expected value is within tolerance.
// Write the theoretical color as the expected value (255, 0, 0 for red). The tolerance
// absorbs the fraction dropped in sampling and the differing rounding policies.
// Coefficient precision is not the reason
static bool CheckPixel(
BYTE y, BYTE u, BYTE v,
MFVideoTransferMatrix matrix,
int expectedR, int expectedG, int expectedB,
int tolerance = 2)
{
BYTE bgra[4] = {};
if (FAILED(ConvertLimitedYuvPixelToBgra(y, u, v, matrix, bgra)))
{
return false;
}
return std::abs(static_cast<int>(bgra[2]) - expectedR) <= tolerance
&& std::abs(static_cast<int>(bgra[1]) - expectedG) <= tolerance
&& std::abs(static_cast<int>(bgra[0]) - expectedB) <= tolerance
&& bgra[3] == 255; // alpha must always be opaque
}
// Usage (BT.601 limited range)
// CheckPixel(16, 128, 128, MFVideoTransferMatrix_BT601, 0, 0, 0); // black
// CheckPixel(235, 128, 128, MFVideoTransferMatrix_BT601, 255, 255, 255); // white
// CheckPixel(81, 90, 240, MFVideoTransferMatrix_BT601, 255, 0, 0); // red (R=254 through the formula)
// CheckPixel(41, 240, 110, MFVideoTransferMatrix_BT601, 0, 0, 255); // blue
The same works for BT.709. The coefficients differ, so the Y/U/V values differ too: red in BT.709, for example, is Y=63, U=102, V=240. Feeding the 601 values straight into the 709 branch drifts the colors, so keeping separate test rows lets you catch a mixed-up matrix on the spot.
In the GitHub sample, this single-pixel conversion is factored out into an OS-independent header, so the test runs even without Windows.
Stage 2: Compare the Pattern A and Pattern B Output
Once a single pixel checks out, move on to the whole frame. From the same time in the same video,
- Pattern A (
MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING+RGB32) - Pattern B (receive
NV12/YUY2and convert it yourself)
take one frame from each of the two paths and compare them pixel by pixel.
How to read the difference:
Take |A.R - B.R|, |A.G - B.G|, |A.B - B.B| for each pixel
Report the maximum, and the percentage of pixels above a threshold
Do not expect an exact match here. There are two reasons.
- The Source Reader’s video processing may be performing the chroma upsampling by something other than nearest neighbor. The manual implementation in 5.7. is a minimal one that uses the shared chroma directly for all 4 pixels, so differences show up most at edges
- Rounding and intermediate precision are handled differently
So what to look at is not “does it match” but the pattern in how the differences appear.
| Difference you see | What to suspect |
|---|---|
| Flat areas match, differences appear only at color boundaries | A difference in chroma upsampling. Expected |
| Everything is uniformly off | The matrix (601 / 709) or the range (16..235 / 0..255) is mixed up |
| Banding, or a diagonal skew | A hard-coded stride. See 7.2. and 7.5. |
| Red and blue are swapped | BGRA and RGBA are mixed up |
| Everything looks transparent / pure black | The 4th byte is not filled with 0xFF. See 7.1. |
Looking at the “shape” of the difference narrows down where to look quite a lot. If the whole image is uniformly off, it is the formula or the color information; if it is localized, it is an index or the stride.
flowchart TB
accTitle: The two-stage verification
accDescr: Diagram showing the two-stage verification - first pass a single known Y/U/V pixel and compare it against hand calculation, then compare the Pattern A and Pattern B frames taken from the same time in the same video and narrow down what to suspect from the shape of the difference.
st5["Stage 1: compare one pixel against hand calculation"] --> st6["Stage 2: compare frames from the two paths"]
st6 --> shp1["Narrow down what to suspect from the shape of the difference"]
st5 -.-> osf1["Needs neither a video nor Media Foundation"]
Figure 19: Clear the single-pixel check first, then compare whole frames, and the shape of the difference points at the cause.
6. Which One Should You Choose?
When in doubt, the following table sorts things out quite well.
| Aspect | Automatic conversion (MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING) |
Manual conversion |
|---|---|---|
| Implementation speed | Excellent | Fair |
| Extracting a few stills | Excellent | Good |
| High frame volume / real-time | Fair | Excellent |
| Explicit control of matrix / range | Fair | Excellent |
| Combining with GPU / D3D | Fair | Good to excellent |
Output formats other than RGB32 |
Fair | Excellent |
| Understanding the fundamentals | Good | Excellent |
For your first implementation, this framing makes it easy.
- Want it working first -> automatic conversion
- Want to own color and performance -> manual conversion
In practice, the sequence “first confirm a correct picture with automatic conversion, then replace it with the manual path” is also quite effective. If you take on everything from the start, it becomes hard to tell where the picture broke.
flowchart TB
accTitle: An order of work that pays off in practice
accDescr: Diagram showing the practical order of work - confirm a correct picture with the automatic conversion first and then replace it with your own manual path, which makes it easier to tell where the picture broke.
step1["Confirm a correct picture with automatic conversion first"] --> step2["Then replace it with the manual path"]
step2 --> good1["Easier to isolate where it broke"]
allin1["Take on everything from the start"] -.-> lost1["Hard to tell where it broke"]
Figure 20: Securing a known-good picture before moving to your own implementation makes it much easier to isolate what broke.
7. Pitfalls That Are Easy to Hit in Practice
7.1. Assuming RGB32 Is RGBA with Alpha
In memory, RGB32 is B, G, R, Alpha or Don't Care.
If you write it out to a PNG as BGRA as is, the 4th byte may be 0, making the image transparent. It is safer to set it to 0xFF before saving.
7.2. Hard-Coding the Stride as width * bytesPerPixel
A very common mistake. The actual sample buffer can contain padding, so the rule is to use the actual stride to move between rows.
7.3. Confusing MF_MT_DEFAULT_STRIDE with the Actual Pitch
MF_MT_DEFAULT_STRIDE is “the minimum stride when that format is represented in contiguous memory.”
For the actual pitch of the sample buffer, prefer the value returned by IMF2DBuffer::Lock2D.
(pitch is another name for stride. As noted in 3.2., this article uses them with the same meaning.)
7.4. Silently Guessing 601 / 709 Without Looking at the Color Metadata
Color bugs are hard to see. They do not crash either. That is what makes them troublesome.
MF_MT_YUV_MATRIXMF_MT_VIDEO_NOMINAL_RANGE
At the very least, look at these. And the right attitude is roughly: values your code does not support should be errors.
7.5. Locating the NV12 UV Plane with width * height
The plane offset is determined by the actual stride and height. Not by width * height.
Do this sloppily and you get shifted colors or corrupted images.
flowchart TB
accTitle: How to locate the NV12 plane boundary
accDescr: Diagram showing that the start of the NV12 UV plane is determined by the actual stride multiplied by the height, and that cutting at width times height leads to color drift and corrupted images.
head1["Start of the buffer (Y plane)"] -->|"Advance stride × height"| uv2["Start of the UV plane"]
wr1["Cut at width × height"] -.-> brk1["Color drift and corrupted images"]
Figure 21: Locate the UV plane boundary with stride x height, never with width x height.
7.6. Processing Interlaced Video Assuming Progressive
The manual samples in this article assume progressive video. Reading interlaced content as if each frame were a single field can produce comb-like artifacts. If you need deinterlacing, it is more natural to consider the Source Reader’s automatic video processing or the Video Processor MFT.
7.7. Ignoring the Quality of 4:2:0 Chroma Upsampling
For clarity, the NV12 conversion in this article uses the shared chroma directly for each pixel. That is sufficient for many uses, but if image quality is the priority, it is worth studying the upconversion approach described in the recommended YUV formats documentation.
8. Summary
When converting YUV to RGB with Media Foundation, keeping the following framework in mind makes it much harder to get lost.
- Behind the decoder,
NV12orYUY2— not RGB — is what normally comes out - If you want the easy path, request
RGB32viaMF_SOURCE_READER_ENABLE_VIDEO_PROCESSING - If you want control, receive
NV12/YUY2and convert to BGRA yourself - On the manual path, get sampling / range / matrix / stride right before worrying about the formula
- Being vague about
BT.601/BT.709,16..235, and4:2:0/4:2:2leads to color drift or broken pictures
YUV -> RGB is a bit unapproachable at first. But once the picture of
NV12shares U/V across 2x2 blocksYUY2shares U/V across 2 horizontal pixels- apply the matrix to that U/V together with Y
settles into your head, it becomes quite tame. Those mysterious cosmic-colored byte sequences start to look like properly meaningful pixels.
flowchart TB
accTitle: The picture to keep in your head
accDescr: Diagram showing that once the picture of NV12 sharing U/V across 2x2 blocks, YUY2 sharing U/V across 2 horizontal pixels, and applying the matrix to that U/V together with Y settles in, a mysterious byte sequence starts to look like meaningful pixels.
im2["NV12: U/V shared across 2x2"] --> ap1["Apply the matrix to that U/V and Y"]
im3["YUY2: U/V shared across 2 horizontal pixels"] --> ap1
ap1 --> see1["The byte sequence starts to look like meaningful pixels"]
Figure 22: Once you hold those two pictures - the sharing unit and the matrix - YUV byte sequences read plainly.
9. References
Sample Code for This Article
Related KomuraSoft Articles
- An Introduction to Media Foundation - Understanding the API Through a COM Lens
- Extracting a Still Image from an MP4 at a Specific Time with Media Foundation
Microsoft Learn
- Source Reader
- Using the Source Reader to Process Media Data
- MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING attribute
- IMFSourceReader::SetCurrentMediaType
- Recommended 8-Bit YUV Formats for Video Rendering
- Extended Color Information
- Uncompressed Video Buffers
- IMF2DBuffer::Lock2D
- MF_MT_VIDEO_NOMINAL_RANGE attribute
- MFVideoTransferMatrix enumeration
- Video Processor MFT
- Uncompressed RGB Video Subtypes
Related Articles
Recent articles sharing the same tags. Deepen your understanding with closely related topics.
How to Burn Images and Text into MP4 Frames with Media Foundation
How to burn an image and text into every frame of an MP4 with Media Foundation and produce a new MP4, organized around the roles of the S...
Extracting a Still Image from an MP4 at a Specific Time with Media Foundation
How to grab the frame closest to a given time in an MP4 with the Source Reader, fix up stride and the RGB32 alpha byte, and save it as a ...
An Introduction to Media Foundation - Understanding the API Through a COM Lens
We explain what Media Foundation is, together with the basic vocabulary of Windows media APIs - COM, HRESULT, IMFSourceReader, MFTs - in ...
Time Travel Debugging — Recording and Rewinding the Bugs That Never Reproduce in Long-Running Apps
A once-a-month bug leaves only its result in a crash dump. Record and rewind execution with WinDbg Time Travel Debugging (TTD): TTD.exe, ...
Why Arguments Break — The Rules of Windows Command-Line Arguments
Windows passes CreateProcess a single string that the receiver splits. Covers the CommandLineToArgvW, CRT, and .NET rules, ArgumentList, ...
Related Topics
These topic pages place the article in a broader service and decision context.
Windows Technical Topics
Topic hub for KomuraSoft LLC's Windows development, investigation, and legacy-asset articles.
Where This Topic Connects
This article connects naturally to the following service pages.
Windows App Development
This topic covers Media Foundation, the Source Reader, image saving, and video frame conversion — a Windows media-processing implementation theme that fits well with our Windows application development service.
Technical Consulting & Design Review
If you want to sort out the division of responsibility for YUV / RGB conversion, color spaces, stride, and conversion-path design up front, this topic works well as a technical consulting / design review engagement.
Frequently Asked Questions
Common questions about the topic of this article.
- Why does a Media Foundation decoder output YUV instead of RGB?
- Because the human eye is more sensitive to fine detail in brightness than in color, video benefits from a design that keeps Y (the brightness-oriented component) at full detail and U/V (the color-difference components) coarser. That is why, in the Windows video world, the uncompressed frames coming out of a decoder are normally YUV-family formats such as NV12 or YUY2. It also helps to read YUV as effectively meaning Y'CbCr in the context of digital video.
- What is the easiest way to get RGB frames?
- Enable MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING on IMFSourceReader and request MFVideoFormat_RGB32. For extracting a few still images or generating thumbnails, this is by far the easiest route. However, this automatic conversion is software processing and is not optimized for real-time playback, so if you need high-volume processing or control over color, receive the frames as YUV and convert them yourself.
- What should I watch out for when converting YUV to RGB myself?
- It is not over once you multiply by three coefficients: subsampling (4:2:0 / 4:2:2), range, matrix, and stride all come into play. The things that most often break colors in practice are not looking at MF_MT_YUV_MATRIX and MF_MT_VIDEO_NOMINAL_RANGE, and assuming the stride is width times bytesPerPixel. The shortest path is to understand the structure of NV12 and YUY2 properly first.
- What is the difference between NV12 and YUY2?
- NV12 is a 4:2:0 format: a Y plane followed by a UV plane in which U and V alternate, and the 4 pixels of a 2x2 block share one U/V pair. YUY2 is a 4:2:2 format in which 2 horizontal pixels share one U/V pair. Both show up often in practice, and what differs is how much the color is thinned out (subsampling).