How to Convert YUV to RGB with Media Foundation

· Updated: · · Media Foundation, C++, Windows Development, Video Processing, YUV

Revision history (1 updates, last updated Sep 1, 2026)

A log of the changes made to this article. Where a pre-update version was archived, it stays readable at a permanent DOI link.

Retranslated as a full translation of the Japanese original. The previous English version was an abridgement that carried only part of the source, so sections, tables, Mermaid diagrams, figure captions and FAQ entries were missing. All of them have been restored to match the Japanese original, and the technical claims are the same as in the Japanese version. Read the version before this update (DOI: 10.5281/zenodo.21614501)
First published
Cite this article(DOI: 10.5281/zenodo.21614500)

This article is archived on Zenodo. Below are both the DOI that always resolves to the latest version and the DOI pinned to the version you are reading.

Go Komura (2026). How to Convert YUV to RGB with Media Foundation. KomuraSoft LLC. https://doi.org/10.5281/zenodo.21614500 https://comcomponent.com/en/blog/2026/03/15/002-media-foundation-yuv-to-rgb-conversion-patterns/

DOI (latest version)
10.5281/zenodo.21614500
DOI (this version)
10.5281/zenodo.22217146

You want to pull a frame out of a video and save it as a PNG, pass it to WIC or GDI, or display it in your UI. In those situations, the application side wants RGB pixel data.

However, the frames that come out of a Media Foundation decoder are, quite commonly, YUV-family formats such as NV12 or YUY2. If you treat the raw byte stream as an image as-is, you get a rather sad picture: broken colors, banding, or a strangely greenish tint.

In an earlier post, An Introduction to Media Foundation - Understanding the API Through a COM Lens, we covered the big picture, and in Extracting a Still Image from an MP4 at a Specific Time with Media Foundation, we covered still-image extraction. This time we tackle the step that sits in between: the YUV -> RGB conversion itself.

In this article, we separate and organize the following two patterns.

  • Pattern A: let IMFSourceReader automatically take the frames all the way to RGB32
  • Pattern B: receive NV12 / YUY2 and convert to RGB yourself

The goal is not to memorize API names. It is to be able to picture, in your head, where in Media Foundation the YUV appears and where it turns into RGB.

The code that appears in this article is published on GitHub as a complete sample set (C++ code for Pattern A / Pattern B, CMake configuration, and tests for the pixel conversion).

media-foundation-yuv-to-rgb-conversion-patterns - komurasoft-blog-samples (GitHub)

Prerequisites for Building and Running

If you are going to carry the code in this article into your own project, this is all you need.

Item Requirement
OS Windows 10 or later
Compiler MSVC from Visual Studio 2019 / 2022 (C++17)
SDK Windows SDK (the Media Foundation headers and import libraries). It is included in the Desktop development with C++ workload in Visual Studio
Build The sample uses CMake 3.20 or later. Creating a Visual Studio project by hand is fine too

There are 4 libraries to link. The article code writes them with #pragma comment(lib, ...), but specifying them in the project settings works exactly the same.

  • mfplat.lib
  • mfreadwrite.lib
  • mfuuid.lib
  • ole32.lib

Also, the code in this article assumes CoInitializeEx and MFStartup have already been called. Only the single-pixel conversion formula (5.6.) is OS-independent, so the GitHub sample factors it out into a separate header that can also be tested with g++ on Linux.

1. The Conclusion First

Summarizing the conclusions up front:

  • For extracting a few still images or generating thumbnails, the easiest route is to enable MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING and request MFVideoFormat_RGB32
  • However, this automatic conversion is software processing and is not optimized for real-time playback
  • If you are going to write your own conversion, the shortest path is to properly understand NV12 and YUY2 first
  • YUV -> RGB is not just “multiply by three coefficients and you are done” — in practice, subsampling, range, matrix, and stride all come into play
  • The Media Foundation documentation broadly uses the term YUV, but for digital video it is easier to read if you assume it effectively means Y’CbCr
  • The things that most often break colors in practice are not looking at MF_MT_YUV_MATRIX and MF_MT_VIDEO_NOMINAL_RANGE, and assuming the stride is width * bytesPerPixel

In short: if you want the easy path, have the Source Reader output RGB32. If you need high-volume processing or control over color, receive the frames as YUV and convert them yourself. Those are the two choices.

The two choices in this articleDiagram showing the two choices this article covers - let the Source Reader output RGB32 if you want the easy path, and receive YUV and convert it yourself if you need high-volume processing or control over color.Take the easy pathHigh volume and color controlYou want RGB framesPattern A: let the Source Reader output RGB32Pattern B: receive YUV and convert it yourself

Figure 1: The choice is Pattern A if you take the easy path, and Pattern B if you take on the throughput and the responsibility for color.

In the diagram a solid line marks a relation that always holds and a dashed line marks a conditional one (the conditions are given per relation on the detail page). The full list of relations (23 in total, with evidence and certainty) and the definitions of the main concepts are collected on the knowledge map detail page (in Japanese). Data: JSON-LD / Turtle

2. Start with a Picture

It is faster to first look at a diagram of what happens inside Media Foundation.

Pattern APattern BMP4 / H.264 / HEVCdecoderYUV frames such as NV12 / YUY2 / YV12Source Reader video processingRGB32Your own conversion codeBGRA / RGB

Figure 2: What the decoder emits is YUV frames, and from there the road to RGB forks into Pattern A and Pattern B.

If the contents of the video file are a compressed format such as H.264 or HEVC, the decoder first turns them back into uncompressed frames. These uncompressed frames are not necessarily RGB. In fact, in the Windows video world, YUV-family formats are the norm.

So when the application wants RGB, you choose one of the following.

  1. Have Media Foundation take the frames all the way to RGB32
  2. Receive YUV and turn it into RGB with your own code

This article is exactly about that fork in the road.

3. Sorting Out the Relationship Between YUV and RGB First

3.1. It Says YUV, but It Is Really About Y’CbCr

Windows API names and documentation broadly use the term YUV. In the context of digital video, however, you can read U as Cb and V as Cr with essentially no problem.

Roughly speaking:

  • Y is the brightness-oriented component
  • U / V are the color-difference components
  • RGB is each pixel directly carrying Red / Green / Blue

That is the relationship.

The human eye is more sensitive to fine detail in brightness than in color. So for video, a design that keeps Y at full detail and U/V somewhat coarser pays off. This is why YUV-family formats are so widely used.

Why YUV-family formats are usedDiagram showing that the human eye is more sensitive to fine detail in brightness than in color, so a design that keeps Y fine and U/V coarse pays off, which is the reason YUV-family formats are widely used.The human eye is sensitive to brightnessKeep Y fine and U/V coarseWhy YUV-family formats are used

Figure 3: The YUV-family design keeps Y fine and U/V coarse to match how sensitive the eye is to brightness.

3.2. 4:4:4 / 4:2:2 / 4:2:0 Is “How Much the Color Is Thinned Out”

This is the key to reading YUV.

Notation Meaning Typical examples
4:4:4 Each pixel has its own Y/U/V AYUV, I444
4:2:2 2 pixels horizontally share one U/V pair YUY2, UYVY, I422
4:2:0 A 2x2 pixel block shares one U/V pair NV12, YV12, I420

It helps a lot to first look at the shape of the two formats you encounter most in practice.

Before that, let us pin down one term. The stride (also called the pitch) is the number of bytes in one row. It is not the image width itself: it is how many bytes you advance to reach the start of the next row, including any padding at the end of the row. This article uses stride and pitch with the same meaning. Microsoft Learn uses both wordings as well, so there is nothing to disambiguate.

In the diagrams below, W is the width, H is the height, and S is the stride. The point is that S >= W, and S == W is not guaranteed.

NV12 (4:2:0, planar) / width = W, height = H, stride = S

  <----------- S bytes ----------->
  <--- W --->
 +-----------+---------------------+  --+
 | Y Y Y Y Y | (padding)           |    |
 | Y Y Y Y Y | (padding)           |    | Y plane
 | Y Y Y Y Y | (padding)           |    | S * H bytes
 | Y Y Y Y Y | (padding)           |    |
 +-----------+---------------------+  --+  <- plane boundary = S * H from the start
 | U V U V U | (padding)           |    |
 | U V U V U | (padding)           |    | UV plane
 +-----------+---------------------+  --+  height is H / 2 rows

  Y of row y      : yPlane  + S * y
  UV of row y     : uvPlane + S * (y / 2)
  Start of UV     : scanline0 + S * H

In NV12, the 4 pixels of a 2x2 block share a single U/V pair. Y exists for each individual pixel. The UV plane uses the same stride as the Y plane, but it has half as many rows. That is why the plane boundary is S * H and not W * H (we come back to this in 7.5.).

YUY2 (4:2:2, packed) / width = W, height = H, stride = S

  <-------------- S bytes --------------->
  <------- W * 2 bytes ------->
 +-----------------------------+----------+
 | Y0 U0 Y1 V0  Y2 U2 Y3 V2 …  | (padding)|   row 0
 | Y0 U0 Y1 V0  Y2 U2 Y3 V2 …  | (padding)|   row 1
 +-----------------------------+----------+

  Start of row y : scanline0 + S * y
  2 pixels = 4 bytes (Y, U, Y, V)
  Only one plane (packed, so there is no boundary)

In YUY2, 2 horizontal pixels share one U/V pair. Y0 and Y1 are separate, but U0 and V0 are shared. Being packed, there is no plane boundary to compute, but moving between rows still uses the stride.

At this point you can already see that YUV -> RGB is not a simple one-pixel-to-one-pixel substitution. First you have to think about how to assign the shared U/V to which pixels.

NV12 and YUY2 sharing units comparedDiagram showing that NV12 shares one U/V pair across the 4 pixels of a 2x2 block while YUY2 shares one pair across 2 horizontal pixels, so you first have to work out how to assign the shared U/V to each pixel.NV12 (4:2:0)4 pixels in a 2x2 block share one U/V pairYUY2 (4:2:2)2 horizontal pixels share one U/V pairDecide how to assign them to each pixel

Figure 4: Both formats share U/V, so the conversion is not a matter of substituting one pixel at a time.

3.3. YUV -> RGB Is “Color Space Conversion + Sampling Conversion”

If you look at Media Foundation’s Extended Color Information, strictly correct color conversion has quite a few stages: inverse quantization, chroma upsampling, YUV -> RGB, the transfer function, primaries conversion, and finally quantization.

That said, as practical code for 8-bit SDR, it is easiest to understand if you split it into the following three layers.

  1. Undo the subsampling Expand the 4:2:0 or 4:2:2 U/V so that every pixel can reference a value
  2. Undo the range Video Y normally uses 16..235 and U/V use 16..240, so undo that scaling
  3. Apply the matrix Convert to RGB using coefficients such as BT.601 or BT.709

In other words, in practical terms, YUV -> RGB conversion is the process of deciding:

  • which U/V is the color for that pixel
  • which coefficients to use to turn that Y/U/V back into RGB
The three layers to get right in practical codeDiagram showing that in practical 8-bit SDR code it is easier to understand YUV to RGB conversion as three layers - undo the subsampling, undo the range, and apply the matrix.YUV frameUndo the subsamplingUndo the range (16..235 and so on)Apply the matrix (601 / 709)RGB

Figure 5: The conversion is not three coefficients; it is built from three layers - subsampling, range, and matrix.

3.4. Treat BT.601 and BT.709 Carelessly and Colors Drift Subtly

The Media Foundation documentation describes the relationship as BT.601 being preferred for SDTV and below, and BT.709 for video beyond SD.

However, silently guessing “the resolution is large, so it must be 709” is not a great idea. Color drift does not crash, so it easily slips into production unnoticed.

Media Foundation can carry color space information as media type attributes. At minimum, look at these two:

  • MF_MT_YUV_MATRIX
  • MF_MT_VIDEO_NOMINAL_RANGE

Looking at these two and explicitly accepting only the combinations your code supports makes it far less likely that something breaks silently later.

A flow that avoids guessing the color spaceDiagram showing that silently guessing 601 or 709 from the resolution lets color drift slip into production unnoticed, so the flow is to look at MF_MT_YUV_MATRIX and MF_MT_VIDEO_NOMINAL_RANGE and explicitly accept only the combinations you support.Silently guess from the resolutionColor drift does not crashIt slips into production unnoticedLook at the matrix and range attributesAccept only the combinations you supportSilent breakage is prevented

Figure 6: Do not guess the color space - read the attributes and explicitly accept only the combinations you can handle.

3.5. The First Formula to Learn Is the BT.601 Limited-Range Version

The canonical 8-bit BT.601 formula looks like this.

C = Y - 16
D = U - 128
E = V - 128

R = clip(1.164383 * C + 1.596027 * E)
G = clip(1.164383 * C - 0.391762 * D - 0.812968 * E)
B = clip(1.164383 * C + 2.017232 * D)

For BT.709 the coefficients change. We will show that in code later.

What matters here is not memorizing the coefficients but the structure: subtract the black level 16 from Y, and view U/V as centered on 128.

The structure of the conversion formulaDiagram showing the structure of the conversion formula - subtract the black level of 16 from Y, view U and V as centered on 128, multiply by the coefficients for each matrix, and clip the result.Subtract 16 from Y (black level)Multiply by the matrix coefficientsSubtract 128 from U / V (center)Clip to 0..255

Figure 7: What you memorize is not the coefficients but the structure of the formula - black level 16 and center 128.

4. Pattern A: Let Media Foundation Convert Automatically

4.1. When This Is a Good Fit

This approach is well suited to situations like the following.

  • You want to extract a single still image from an MP4
  • You want to create a few thumbnails
  • You want an RGB image to hand to WIC
  • Batch or tooling use is fine; this is not real-time playback

The Source Reader has a feature that performs limited video processing of YUV -> RGB32 when you use MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING.

However, as Microsoft Learn also notes, this is software processing and not optimized for playback. If you want to process hundreds of frames per second, leaning on this is not quite the right tool.

What the automatic conversion suits and does not suitDiagram showing that the Source Reader automatic conversion is software processing that is not optimized for playback, so it suits still-image extraction, thumbnails, and batch work, while leaning on it for hundreds of frames per second is the wrong tool.Good fitDo not lean on itSource Reader automatic conversionSoftware processingStills, thumbnails, batch workHundreds of frames per second in real time

Figure 8: The automatic conversion is software processing, so keep it to tooling that handles a small number of frames.

4.2. What You Set to Get RGB32 Out

The flow is quite straightforward.

  1. In the attributes passed to MFCreateSourceReaderFromURL, set MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING = TRUE
  2. Select the video stream
  3. Request MFMediaType_Video / MFVideoFormat_RGB32 via SetCurrentMediaType
  4. Read samples with ReadSample

That alone makes the limited video processing inserted behind the decoder do the YUV -> RGB32 for you.

The four steps that turn on automatic conversionDiagram showing the four steps of automatic conversion - enable video processing in the attributes and create the reader, select the video stream, request RGB32, and read with ReadSample.Enable video processing in the attributesSelect the video streamRequest RGB32Read with ReadSampleConverted to RGB32 behind the decoder

Figure 9: Follow those four steps and the video processing behind the decoder carries the frames all the way to RGB32.

4.3. Code

The following code assumes CoInitializeEx and MFStartup have already been done. A minimal version looks roughly like this.

#include <windows.h>
#include <mfapi.h>
#include <mfidl.h>
#include <mfreadwrite.h>
#include <mferror.h>
#include <wrl/client.h>

#pragma comment(lib, "mfplat.lib")
#pragma comment(lib, "mfreadwrite.lib")
#pragma comment(lib, "mfuuid.lib")
#pragma comment(lib, "ole32.lib")

using Microsoft::WRL::ComPtr;

HRESULT CreateSourceReaderWithAutoRgb(
    const wchar_t* path,
    IMFSourceReader** ppReader)
{
    if (!path || !ppReader) return E_POINTER;
    *ppReader = nullptr;

    ComPtr<IMFAttributes> attrs;
    HRESULT hr = MFCreateAttributes(&attrs, 2);
    if (FAILED(hr)) return hr;

    hr = attrs->SetUINT32(MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING, TRUE);
    if (FAILED(hr)) return hr;

    hr = MFCreateSourceReaderFromURL(path, attrs.Get(), ppReader);
    if (FAILED(hr)) return hr;

    hr = (*ppReader)->SetStreamSelection(MF_SOURCE_READER_ALL_STREAMS, FALSE);
    if (FAILED(hr)) return hr;

    hr = (*ppReader)->SetStreamSelection(MF_SOURCE_READER_FIRST_VIDEO_STREAM, TRUE);
    if (FAILED(hr)) return hr;

    ComPtr<IMFMediaType> outType;
    hr = MFCreateMediaType(&outType);
    if (FAILED(hr)) return hr;

    hr = outType->SetGUID(MF_MT_MAJOR_TYPE, MFMediaType_Video);
    if (FAILED(hr)) return hr;

    hr = outType->SetGUID(MF_MT_SUBTYPE, MFVideoFormat_RGB32);
    if (FAILED(hr)) return hr;

    hr = (*ppReader)->SetCurrentMediaType(
        MF_SOURCE_READER_FIRST_VIDEO_STREAM,
        nullptr,
        outType.Get());
    if (FAILED(hr)) return hr;

    return S_OK;
}

HRESULT ReadOneRgb32Sample(
    IMFSourceReader* reader,
    IMFSample** ppSample,
    LONGLONG* pTimestamp100ns)
{
    if (!reader || !ppSample) return E_POINTER;
    *ppSample = nullptr;
    if (pTimestamp100ns) *pTimestamp100ns = 0;

    DWORD streamIndex = 0;
    DWORD flags = 0;
    LONGLONG timestamp = 0;

    HRESULT hr = reader->ReadSample(
        MF_SOURCE_READER_FIRST_VIDEO_STREAM,
        0,
        &streamIndex,
        &flags,
        &timestamp,
        ppSample);

    if (FAILED(hr)) return hr;
    if (flags & MF_SOURCE_READERF_ENDOFSTREAM) return MF_E_END_OF_STREAM;
    if (*ppSample == nullptr) return MF_E_INVALID_STREAM_DATA;

    if (pTimestamp100ns) *pTimestamp100ns = timestamp;
    return S_OK;
}

After this, calling GetCurrentMediaType lets you check the actual output size and stride.

4.4. Strengths of This Approach

The good thing about this approach is that it gets you to a correct picture quickly.

  • You do not have to write the 4:2:0 / 4:2:2 expansion yourself
  • It hides much of the hassle of matrix handling / deinterlacing
  • The output is easy to hand to WIC or GDI
  • For processing a handful of frames, it is perfectly practical

For still-image extraction tools, starting here is quite natural.

4.5. But There Are Pitfalls Too

This automatic conversion has the following characteristics.

Item Details
Conversion target Basically RGB32
Implementation Software processing
Suited for Small numbers of frames, thumbnails, offline processing
Not suited for D3D-based real-time rendering, high-volume frame processing
Incompatible attributes MF_SOURCE_READER_D3D_MANAGER, MF_READWRITE_DISABLE_CONVERTERS

And one more important thing: the handling of the 4th byte in RGB32. In memory, Windows RGB32 is laid out as Blue / Green / Red / Alpha or Don’t Care. It is not ARGB32. If you pass it to WIC as 32bppBGRA, it is safer to fill the 4th byte with 0xFF to make it opaque.

Handling the 4th byte of RGB32Diagram showing that the 4th byte after B, G, and R in Windows RGB32 is either alpha or don't care, so filling it with 0xFF to make it opaque before handing the data to WIC as 32bppBGRA is safer.Fill it with 0xFFHand it over as isRGB32 memory layout3 bytes of B, G, RThe 4th byte is alpha or don't careCan be handed to WIC as 32bppBGRAIt can come out transparent

Figure 10: The 4th byte is undefined, so make it opaque with 0xFF before handing the data to WIC.

We touched on this as an easy thing to trip over in the previous still-image extraction article as well.

5. Pattern B: Write the Conversion Yourself

5.1. When This Is a Good Fit

Doing the conversion yourself is a good fit in cases like these.

  • You process a large number of frames and want to optimize the conversion yourself
  • You want to feed NV12 straight to the GPU or SIMD code
  • You want to handle BT.601 / BT.709 / range explicitly
  • You want to produce output formats other than RGB32
  • The Source Reader’s limited automatic conversion is not enough

You could call it the pattern where you take on responsibility for throughput and color in exchange for freedom.

The trade-off of manual conversionDiagram showing that manual conversion takes on responsibility for throughput and color in exchange for optimization, a path into GPU and SIMD code, explicit control of matrix and range, and freedom in the output format.Choose manual conversionYou own throughput and colorFreedom in optimization, GPU / SIMD, and output formatExplicit control of matrix / range

Figure 11: Manual conversion trades responsibility for performance and freedom over color.

5.2. Overall Flow of Manual Conversion

The steps are as follows.

  1. Set the Source Reader output to NV12 or YUY2
  2. Get the actual subtype and attributes with GetCurrentMediaType
  3. Check MF_MT_FRAME_SIZE, MF_MT_DEFAULT_STRIDE, MF_MT_YUV_MATRIX, and MF_MT_VIDEO_NOMINAL_RANGE
  4. Take the buffer out of the sample and lock it
  5. Work out the Y/U/V that each pixel references
  6. Apply the matrix and write BGRA

The code in this article narrows the scope to 8-bit SDR / progressive / NV12 or YUY2 / limited range. Narrowing the assumptions here is not laziness; it actually matters. A YUV converter written to “accept everything for now” tends to break colors silently.

Overall flow of manual conversionDiagram showing the manual conversion steps - request NV12 or YUY2, check the actual media type and attributes, lock the buffer, work out the Y/U/V for each pixel, and apply the matrix to write BGRA.Request NV12 / YUY2Check the actual subtype and attributesLock the bufferWork out the Y/U/V for each pixelApply the matrix and write BGRANarrowing the assumptions keeps colors from breaking

Figure 12: Manual conversion runs in the order request, check, lock, reference, convert, and the narrower the assumptions the safer it gets.

5.3. First, Specify the Output Media Type Explicitly

First, tell the Source Reader that you want the YUV as-is. This also assumes CoInitializeEx / MFStartup have already been done.

#include <windows.h>
#include <mfapi.h>
#include <mfidl.h>
#include <mfreadwrite.h>
#include <mferror.h>
#include <wrl/client.h>

using Microsoft::WRL::ComPtr;

HRESULT ConfigureSourceReaderForSubtype(
    IMFSourceReader* reader,
    REFGUID subtype)
{
    if (!reader) return E_POINTER;

    HRESULT hr = reader->SetStreamSelection(MF_SOURCE_READER_ALL_STREAMS, FALSE);
    if (FAILED(hr)) return hr;

    hr = reader->SetStreamSelection(MF_SOURCE_READER_FIRST_VIDEO_STREAM, TRUE);
    if (FAILED(hr)) return hr;

    ComPtr<IMFMediaType> outType;
    hr = MFCreateMediaType(&outType);
    if (FAILED(hr)) return hr;

    hr = outType->SetGUID(MF_MT_MAJOR_TYPE, MFMediaType_Video);
    if (FAILED(hr)) return hr;

    hr = outType->SetGUID(MF_MT_SUBTYPE, subtype);
    if (FAILED(hr)) return hr;

    hr = reader->SetCurrentMediaType(
        MF_SOURCE_READER_FIRST_VIDEO_STREAM,
        nullptr,
        outType.Get());
    if (FAILED(hr)) return hr;

    return S_OK;
}

Here you pass either MFVideoFormat_NV12 or MFVideoFormat_YUY2 as subtype.

What you need to watch out for is that the subtype you requested does not necessarily go through. Confirm what actually comes out with GetCurrentMediaType.

Separating what you request from what actually comes outDiagram showing that the subtype requested through SetCurrentMediaType does not necessarily go through, so the flow is to confirm what actually comes out with GetCurrentMediaType.May not go through as isRequest a subtypeThe actual outputConfirm it with GetCurrentMediaTypeWrite the rest of the code against the confirmed values

Figure 13: A request is only a request - always confirm the actual output with GetCurrentMediaType before you use it.

5.4. Before Converting, Accept Only the Color Information You Support

For a manual conversion, first pull the minimum information from the media type. The sample in this article accepts only NV12 / YUY2, and lets through only BT.601 or BT.709 for the matrix and only MFNominalRange_16_235 for the range.

#include <vector>

struct DecodedFrameInfo
{
    GUID subtype = GUID_NULL;
    UINT32 width = 0;
    UINT32 height = 0;
    LONG defaultStride = 0;
    MFVideoTransferMatrix matrix = MFVideoTransferMatrix_Unknown;
    MFNominalRange nominalRange = MFNominalRange_Unknown;
};

HRESULT GetDefaultStride(
    IMFMediaType* pType,
    LONG* plStride)
{
    if (!pType || !plStride) return E_POINTER;

    LONG stride = 0;
    HRESULT hr = pType->GetUINT32(
        MF_MT_DEFAULT_STRIDE,
        reinterpret_cast<UINT32*>(&stride));

    if (FAILED(hr))
    {
        GUID subtype = GUID_NULL;
        UINT32 width = 0;
        UINT32 height = 0;

        hr = pType->GetGUID(MF_MT_SUBTYPE, &subtype);
        if (FAILED(hr)) return hr;

        hr = MFGetAttributeSize(pType, MF_MT_FRAME_SIZE, &width, &height);
        if (FAILED(hr)) return hr;

        hr = MFGetStrideForBitmapInfoHeader(subtype.Data1, width, &stride);
        if (FAILED(hr)) return hr;

        hr = pType->SetUINT32(MF_MT_DEFAULT_STRIDE, static_cast<UINT32>(stride));
        if (FAILED(hr)) return hr;
    }

    *plStride = stride;
    return S_OK;
}

HRESULT GetStrictDecodedFrameInfo(
    IMFMediaType* pType,
    DecodedFrameInfo* pInfo)
{
    if (!pType || !pInfo) return E_POINTER;

    HRESULT hr = pType->GetGUID(MF_MT_SUBTYPE, &pInfo->subtype);
    if (FAILED(hr)) return hr;

    if (pInfo->subtype != MFVideoFormat_NV12 &&
        pInfo->subtype != MFVideoFormat_YUY2)
    {
        return MF_E_INVALIDMEDIATYPE;
    }

    hr = MFGetAttributeSize(pType, MF_MT_FRAME_SIZE, &pInfo->width, &pInfo->height);
    if (FAILED(hr)) return hr;

    hr = GetDefaultStride(pType, &pInfo->defaultStride);
    if (FAILED(hr)) return hr;

    UINT32 value = 0;

    hr = pType->GetUINT32(MF_MT_YUV_MATRIX, &value);
    if (FAILED(hr)) return hr;

    pInfo->matrix = static_cast<MFVideoTransferMatrix>(value);
    if (pInfo->matrix != MFVideoTransferMatrix_BT601 &&
        pInfo->matrix != MFVideoTransferMatrix_BT709)
    {
        return MF_E_INVALIDMEDIATYPE;
    }

    hr = pType->GetUINT32(MF_MT_VIDEO_NOMINAL_RANGE, &value);
    if (FAILED(hr)) return hr;

    pInfo->nominalRange = static_cast<MFNominalRange>(value);
    if (pInfo->nominalRange != MFNominalRange_16_235)
    {
        return MF_E_INVALIDMEDIATYPE;
    }

    return S_OK;
}

This is deliberately strict. The Media Foundation enum documentation does say things like “treat Unknown as BT.709,” but in practice, silently rounding here makes color drift hard to notice. At least in a first implementation, returning an error for combinations you do not support is safer.

When Does Unknown Come Back?

In case this feels too strict, here are the paths that produce Unknown. They are mostly cases where the source video simply does not carry color information.

  • The H.264 / HEVC VUI has no color description. Per the specification, when colour_description_present_flag is 0, matrix_coefficients is treated as unspecified. If that missing information goes through the decoder, the matrix passed downstream is unspecified as well
  • Raw YUV from a capture device or an old container. These are paths that carry no color space description at all
  • Sometimes the MF_MT_YUV_MATRIX attribute is not present at all. In that case GetUINT32 returns no value and fails with MF_E_ATTRIBUTENOTFOUND (the code above rejects it right there via FAILED(hr))

The important point is that Unknown does not mean “we know it is BT.709”; it means “we do not know.” Applying 709 to SD material drifts the colors, and so does the reverse.

From there, the approach splits in two.

  • Reject it strictly (the approach in this article): return an error as unsupported and let the layer above decide that this material is not handled. Saying outright that you cannot handle it is safer than letting colors drift quietly
  • Pick a default and let it through: if you really have to let it through, write down in the log what you assumed when the value was Unknown. Then state explicitly that you picked 601 / 709 based on the resolution

Either way, the one thing to avoid is silently rounding it off. Color drift does not crash, so it rides into production unnoticed.

What to do when the matrix is UnknownDiagram showing that Unknown means you do not know rather than knowing it is BT.709, so the approach splits into rejecting it strictly with an error or letting it through with the assumption written to the log, and that silently rounding it off is the one thing to avoid.The approach in this articleIf you have to let it throughThe one thing to avoidmatrix is Unknown = we do not knowReject it strictly with an errorLog what you assumed and let it throughRound it off silently

Figure 14: Unknown means unknown, so either reject it or let it through with a log entry - never round it off silently.

With cameras and JPEG-family sources, you sometimes want to handle full-range paths separately. Rather than quietly serving both here, the approach is to explicitly narrow the assumptions this code accepts.

5.5. Read the Buffer Trusting the Stride

This part is quite important too.

  • MF_MT_DEFAULT_STRIDE is the minimum stride
  • The actual sample buffer may have an actual stride that includes padding
  • If IMF2DBuffer::Lock2D is available, prefer it

Taking the helper pattern from Microsoft Learn’s Uncompressed Video Buffers and making it directly usable gives us this.

class BufferLock
{
public:
    explicit BufferLock(IMFMediaBuffer* buffer)
        : m_buffer(buffer),
          m_2dBuffer(nullptr),
          m_locked(false)
    {
        if (m_buffer)
        {
            m_buffer->AddRef();
            m_buffer->QueryInterface(IID_PPV_ARGS(&m_2dBuffer));
        }
    }

    ~BufferLock()
    {
        Unlock();

        if (m_2dBuffer)
        {
            m_2dBuffer->Release();
            m_2dBuffer = nullptr;
        }

        if (m_buffer)
        {
            m_buffer->Release();
            m_buffer = nullptr;
        }
    }

    HRESULT Lock(
        LONG defaultStride,
        DWORD heightInPixels,
        BYTE** ppScanline0,
        LONG* pActualStride)
    {
        if (!m_buffer || !ppScanline0 || !pActualStride) return E_POINTER;
        if (m_locked) return MF_E_INVALIDREQUEST;

        if (m_2dBuffer)
        {
            HRESULT hr = m_2dBuffer->Lock2D(ppScanline0, pActualStride);
            if (FAILED(hr)) return hr;

            m_locked = true;
            return S_OK;
        }

        BYTE* pData = nullptr;
        HRESULT hr = m_buffer->Lock(&pData, nullptr, nullptr);
        if (FAILED(hr)) return hr;

        *pActualStride = defaultStride;
        if (defaultStride < 0)
        {
            *ppScanline0 =
                pData + static_cast<size_t>(-defaultStride) * (heightInPixels - 1);
        }
        else
        {
            *ppScanline0 = pData;
        }

        m_locked = true;
        return S_OK;
    }

    void Unlock()
    {
        if (!m_locked) return;

        if (m_2dBuffer)
        {
            m_2dBuffer->Unlock2D();
        }
        else
        {
            m_buffer->Unlock();
        }

        m_locked = false;
    }

private:
    IMFMediaBuffer* m_buffer;
    IMF2DBuffer* m_2dBuffer;
    bool m_locked;
};

The recommended YUV surface definitions use top-left origin and a positive stride, but for actual buffer access it is safer to use the stride (that is, the pitch) the API handed back. Hard-code a width-based value here and things break silently later.

Stride priority orderDiagram showing the priority order for stride - MF_MT_DEFAULT_STRIDE is the minimum stride, the actual buffer may carry padding, so prefer the value Lock2D returns when it is available and avoid hard-coding a width-based value.First choice when availablefallbackBreaks silentlyActual stride returned by Lock2DThe value used for buffer accessMF_MT_DEFAULT_STRIDE (the minimum)Hard-coded width-based value

Figure 15: For moving between rows, prefer the stride Lock2D actually reports and never derive it from the width.

5.6. Turning the Per-Pixel Conversion Formula into Code

Here we handle only the limited range of BT.601 and BT.709. The output is BGRA32, which is easy to hand to WIC or GDI.

inline BYTE ClampToByte(double value)
{
    if (value <= 0.0) return 0;
    if (value >= 255.0) return 255;
    return static_cast<BYTE>(value + 0.5);
}

HRESULT ConvertLimitedYuvPixelToBgra(
    BYTE y,
    BYTE u,
    BYTE v,
    MFVideoTransferMatrix matrix,
    BYTE* dstPixel)
{
    if (!dstPixel) return E_POINTER;

    const double c = static_cast<double>(y) - 16.0;
    const double d = static_cast<double>(u) - 128.0;
    const double e = static_cast<double>(v) - 128.0;

    double r = 0.0;
    double g = 0.0;
    double b = 0.0;

    switch (matrix)
    {
    case MFVideoTransferMatrix_BT601:
        r = 1.164383 * c + 1.596027 * e;
        g = 1.164383 * c - 0.391762 * d - 0.812968 * e;
        b = 1.164383 * c + 2.017232 * d;
        break;

    case MFVideoTransferMatrix_BT709:
        r = 1.164383 * c + 1.792741 * e;
        g = 1.164383 * c - 0.213249 * d - 0.532909 * e;
        b = 1.164383 * c + 2.112402 * d;
        break;

    default:
        return MF_E_INVALIDMEDIATYPE;
    }

    dstPixel[0] = ClampToByte(b);
    dstPixel[1] = ClampToByte(g);
    dstPixel[2] = ClampToByte(r);
    dstPixel[3] = 255;

    return S_OK;
}

What this does is simple.

  • Subtract 16 from Y
  • Subtract 128 from U / V
  • Multiply by the coefficients for the given matrix
  • Clip the result to 0..255
  • Set the 4th BGRA byte to 255

5.7. Converting NV12 to BGRA32

NV12 is 4:2:0, so the 4 pixels of a 2x2 block share the same U/V. As a minimal implementation, the most understandable approach is to use that shared chroma directly for all 4 pixels.

HRESULT ConvertNv12ToBgra32(
    IMFMediaBuffer* buffer,
    const DecodedFrameInfo& info,
    std::vector<BYTE>& dstBgra)
{
    if (!buffer) return E_POINTER;
    if (info.subtype != MFVideoFormat_NV12) return MF_E_INVALIDMEDIATYPE;
    if ((info.width & 1u) != 0 || (info.height & 1u) != 0)
    {
        return MF_E_INVALIDMEDIATYPE;
    }

    dstBgra.resize(static_cast<size_t>(info.width) * info.height * 4);

    BufferLock lock(buffer);

    BYTE* scanline0 = nullptr;
    LONG actualStride = 0;
    HRESULT hr = lock.Lock(
        info.defaultStride,
        info.height,
        &scanline0,
        &actualStride);
    if (FAILED(hr)) return hr;

    if (actualStride <= 0)
    {
        lock.Unlock();
        return MF_E_INVALIDMEDIATYPE;
    }

    const BYTE* yPlane = scanline0;

    // The UV plane starts stride * height bytes into the buffer.
    // Note that it is not width * height (see the diagram in 3.2.)
    const BYTE* uvPlane =
        scanline0 + static_cast<size_t>(actualStride) * info.height;

    for (UINT32 y = 0; y < info.height; ++y)
    {
        // Always move between rows in units of the stride
        const BYTE* yRow = yPlane + static_cast<size_t>(actualStride) * y;

        // 4:2:0, so 2 vertical rows share one UV row -> y / 2
        // The UV plane uses the same stride as the Y plane
        const BYTE* uvRow = uvPlane + static_cast<size_t>(actualStride) * (y / 2);

        // The output is tightly packed BGRA with no padding, so width * 4
        BYTE* dstRow =
            dstBgra.data() + static_cast<size_t>(info.width) * 4 * y;

        for (UINT32 x = 0; x < info.width; ++x)
        {
            const BYTE Y = yRow[x];

            // The UV plane alternates [U, V].
            // 2 horizontal pixels share one pair, so (x / 2) gives the pair index,
            // and since one pair is 2 bytes, * 2 turns it into a byte offset. +0 is U, +1 is V.
            //   x = 0, 1 -> uvRow[0], uvRow[1]
            //   x = 2, 3 -> uvRow[2], uvRow[3]
            const BYTE U = uvRow[(x / 2) * 2 + 0];
            const BYTE V = uvRow[(x / 2) * 2 + 1];

            hr = ConvertLimitedYuvPixelToBgra(
                Y,
                U,
                V,
                info.matrix,
                dstRow + static_cast<size_t>(x) * 4);
            if (FAILED(hr))
            {
                lock.Unlock();
                return hr;
            }
        }
    }

    lock.Unlock();
    return S_OK;
}

This code interprets the chroma upsampling in a nearest-neighbor fashion. Visually that is often perfectly serviceable, but if you are aiming for maximum quality, a design that first performs the 4:2:0 -> 4:2:2 -> 4:4:4 upconversion, as described in the Microsoft Learn YUV article, is theoretically cleaner.

Two designs for chroma upsamplingDiagram showing that the minimal implementation reuses the shared chroma for all 4 pixels in a nearest-neighbor fashion and is often perfectly serviceable, while upconverting from 4:2:0 through 4:2:2 to 4:4:4 before converting is theoretically cleaner when image quality comes first.Minimal: use the shared chroma as isOften perfectly serviceable visuallyUpconvert first, then convertTheoretically cleaner, quality firstIn the order 4:2:0 → 4:2:2 → 4:4:4

Figure 16: Choosing between the minimal implementation that reuses the shared chroma and a design that upconverts in stages.

5.8. Converting YUY2 to BGRA32

YUY2 is packed 4:2:2. Two pixels simply share one U/V pair, so it is a bit easier to read than NV12.

#include <cstddef>

HRESULT ConvertYuy2ToBgra32(
    IMFMediaBuffer* buffer,
    const DecodedFrameInfo& info,
    std::vector<BYTE>& dstBgra)
{
    if (!buffer) return E_POINTER;
    if (info.subtype != MFVideoFormat_YUY2) return MF_E_INVALIDMEDIATYPE;
    if ((info.width & 1u) != 0) return MF_E_INVALIDMEDIATYPE;

    dstBgra.resize(static_cast<size_t>(info.width) * info.height * 4);

    BufferLock lock(buffer);

    BYTE* scanline0 = nullptr;
    LONG actualStride = 0;
    HRESULT hr = lock.Lock(
        info.defaultStride,
        info.height,
        &scanline0,
        &actualStride);
    if (FAILED(hr)) return hr;

    for (UINT32 y = 0; y < info.height; ++y)
    {
        const BYTE* src =
            scanline0 +
            static_cast<ptrdiff_t>(actualStride) * static_cast<ptrdiff_t>(y);

        BYTE* dstRow =
            dstBgra.data() + static_cast<size_t>(info.width) * 4 * y;

        for (UINT32 x = 0; x < info.width; x += 2)
        {
            const BYTE Y0 = src[0];
            const BYTE U  = src[1];
            const BYTE Y1 = src[2];
            const BYTE V  = src[3];

            hr = ConvertLimitedYuvPixelToBgra(
                Y0,
                U,
                V,
                info.matrix,
                dstRow + static_cast<size_t>(x) * 4);
            if (FAILED(hr))
            {
                lock.Unlock();
                return hr;
            }

            hr = ConvertLimitedYuvPixelToBgra(
                Y1,
                U,
                V,
                info.matrix,
                dstRow + static_cast<size_t>(x + 1) * 4);
            if (FAILED(hr))
            {
                lock.Unlock();
                return hr;
            }

            src += 4;
        }
    }

    lock.Unlock();
    return S_OK;
}

Because YUY2 lays the bytes out as Y0 U Y1 V, the structure of “reuse the U/V for every 2 pixels” is visible directly in the data. That makes the mental model easier to build than for NV12.

5.9. The Entry Point When Calling from a Sample

Finally, pulling a contiguous buffer out of the IMFSample and branching on the subtype makes this easy to use.

HRESULT ConvertSampleToBgra32(
    IMFSample* sample,
    const DecodedFrameInfo& info,
    std::vector<BYTE>& dstBgra)
{
    if (!sample) return E_POINTER;

    ComPtr<IMFMediaBuffer> buffer;
    HRESULT hr = sample->ConvertToContiguousBuffer(&buffer);
    if (FAILED(hr)) return hr;

    if (info.subtype == MFVideoFormat_NV12)
    {
        return ConvertNv12ToBgra32(buffer.Get(), info, dstBgra);
    }

    if (info.subtype == MFVideoFormat_YUY2)
    {
        return ConvertYuy2ToBgra32(buffer.Get(), info, dstBgra);
    }

    return MF_E_INVALIDMEDIATYPE;
}

With that, the earlier stage becomes

  • create the reader
  • request NV12 or YUY2
  • build a DecodedFrameInfo from GetCurrentMediaType
  • ReadSample
  • ConvertSampleToBgra32

as a flow.

Branching in the entry-point functionDiagram showing the structure of the entry-point function - pull a contiguous buffer out of the sample, branch to the NV12 conversion when the subtype is NV12 and to the YUY2 conversion when it is YUY2, and return an error for anything else.NV12YUY2Anything elseIMFSamplePull out a contiguous bufferTo the NV12 conversionTo the YUY2 conversionReturn an error

Figure 17: Make the buffer contiguous at the entry point, branch per subtype, and let anything unsupported fall straight through to an error.

The actual calling side looks something like this.

ComPtr<IMFMediaType> currentType;
HRESULT hr = reader->GetCurrentMediaType(
    MF_SOURCE_READER_FIRST_VIDEO_STREAM,
    &currentType);
if (FAILED(hr)) return hr;

DecodedFrameInfo info;
hr = GetStrictDecodedFrameInfo(currentType.Get(), &info);
if (FAILED(hr)) return hr;

DWORD flags = 0;
LONGLONG timestamp = 0;
ComPtr<IMFSample> sample;

hr = reader->ReadSample(
    MF_SOURCE_READER_FIRST_VIDEO_STREAM,
    0,
    nullptr,
    &flags,
    &timestamp,
    &sample);
if (FAILED(hr)) return hr;
if (flags & MF_SOURCE_READERF_ENDOFSTREAM) return MF_E_END_OF_STREAM;
if (!sample) return MF_E_INVALID_STREAM_DATA;

std::vector<BYTE> bgra;
hr = ConvertSampleToBgra32(sample.Get(), info, bgra);
if (FAILED(hr)) return hr;

// bgra can be treated as top-down / 32bpp BGRA

5.10. Where to Put the “Manual Conversion”

The code so far takes the form of the application converting after the Source Reader. That is the easiest to understand.

However, if you want to insert the conversion inside the Media Foundation pipeline, there are other designs.

  • Write your own MFT
  • Use the Video Processor MFT / XVP
  • Write an NV12 -> RGB shader on the GPU side

Going that far changes the topic somewhat, so this article focused on application-side code. Still, it is useful to know that between “let Media Foundation handle it” and “do everything in the app,” there is a middle ground: the Video Processor MFT.

5.11. Confirming That the Conversion Is Correct

Color bugs are hard to see, so check “it runs” and “it is correct” separately. There are two stages, in this order.

Stage 1: Feed in Known Values and Compare Against Hand Calculation

Rather than starting with a video, it is more reliable to pass known Y/U/V values to ConvertLimitedYuvPixelToBgra. You need neither a video file nor Media Foundation.

For BT.601 limited range, here are the Y/U/V values for representative colors and the values you should expect from the formula in 5.6.

Color Y U V Expected R G B
Black 16 128 128 0 0 0
White 235 128 128 255 255 255
Red 81 90 240 254 0 0
Blue 41 240 110 0 0 255

For red, for example, put C = 81 - 16 = 65, D = 90 - 128 = -38, and E = 240 - 128 = 112 into the formula, and you get

R = 1.164383 * 65 + 1.596027 * 112       = 254.44  -> 254
G = 1.164383 * 65 - 0.391762 * (-38)
                  - 0.812968 * 112       =  -0.48  ->   0
B = 1.164383 * 65 + 2.017232 * (-38)     =  -0.97  ->   0

The output is in BGRA order, so as a byte sequence it is 00 00 FE FF.

What matters here is that red comes out as 254 rather than 255. The reason is not the precision of the coefficients. It is that the input Y/U/V are already rounded integers.

Taking the theoretical red (255, 0, 0) down into BT.601 limited range gives Y = 16 + 219 × 0.299 = 81.481, U = 90.203, and V lands exactly on 240. The moment you store that as an 8-bit sample, the fraction disappears and Y becomes 81. The lost 0.481 turns into a shortfall of 0.481 × 1.164383 ≈ 0.56 on the way back. 255 - 0.56 = 254.44 - that is where the 254.44 above comes from. Even with infinite-precision coefficients the result stays 254.44; rounding to 6 decimal places only shows up below the fourth decimal place and never reaches the 8-bit output.

How you finally convert to integers also affects the result. ClampToByte in 5.6. constrains the value to [0, 255] and then truncates value + 0.5, which is round-half-up. With plain truncation (static_cast<BYTE>(value)), this red is still 254, but values sitting near a boundary, such as blue’s B = 255.04 or red’s R = 0.38, come out 1 off. Before you compare against another implementation, find out which one it uses.

In other words, the reason to allow a difference of 1 or 2 is not “the coefficients have different precision” but the two facts that “sampling dropped a fraction” and “the rounding policy differs between implementations.” Put the other way around, a difference those two cannot explain is a real bug. Red coming out as 250, red and blue swapped, shadows lifted - for differences like these, suspect the assumptions of the conversion (BT.601 mistaken for BT.709, full range mistaken for limited range, U and V swapped, a misread stride) rather than the precision of the coefficients. Write these off as “a precision issue” and you miss bugs you could have fixed.

Drawing the line between tolerance and a real bugDiagram showing that a difference of one or two can be explained by the fraction dropped in sampling and by the differing rounding policies, and that a difference those two cannot explain is a real bug whose conversion assumptions should be suspected.A difference of 1 to 2A difference with no explanationLook at the difference from the expected valueExplained by the sampling fraction and the rounding policyA real bugSuspect the matrix, range, U/V, and stride assumptions

Figure 18: Small differences are explained by sampling and rounding; a difference they cannot explain points to a mistaken assumption.

Written as a test, this shape is enough.

#include <cstdlib>  // std::abs

// Check whether the difference from the expected value is within tolerance.
// Write the theoretical color as the expected value (255, 0, 0 for red). The tolerance
// absorbs the fraction dropped in sampling and the differing rounding policies.
// Coefficient precision is not the reason
static bool CheckPixel(
    BYTE y, BYTE u, BYTE v,
    MFVideoTransferMatrix matrix,
    int expectedR, int expectedG, int expectedB,
    int tolerance = 2)
{
    BYTE bgra[4] = {};
    if (FAILED(ConvertLimitedYuvPixelToBgra(y, u, v, matrix, bgra)))
    {
        return false;
    }

    return std::abs(static_cast<int>(bgra[2]) - expectedR) <= tolerance
        && std::abs(static_cast<int>(bgra[1]) - expectedG) <= tolerance
        && std::abs(static_cast<int>(bgra[0]) - expectedB) <= tolerance
        && bgra[3] == 255;  // alpha must always be opaque
}

// Usage (BT.601 limited range)
// CheckPixel(16, 128, 128, MFVideoTransferMatrix_BT601, 0, 0, 0);      // black
// CheckPixel(235, 128, 128, MFVideoTransferMatrix_BT601, 255, 255, 255); // white
// CheckPixel(81, 90, 240, MFVideoTransferMatrix_BT601, 255, 0, 0);     // red (R=254 through the formula)
// CheckPixel(41, 240, 110, MFVideoTransferMatrix_BT601, 0, 0, 255);    // blue

The same works for BT.709. The coefficients differ, so the Y/U/V values differ too: red in BT.709, for example, is Y=63, U=102, V=240. Feeding the 601 values straight into the 709 branch drifts the colors, so keeping separate test rows lets you catch a mixed-up matrix on the spot.

In the GitHub sample, this single-pixel conversion is factored out into an OS-independent header, so the test runs even without Windows.

Stage 2: Compare the Pattern A and Pattern B Output

Once a single pixel checks out, move on to the whole frame. From the same time in the same video,

  1. Pattern A (MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING + RGB32)
  2. Pattern B (receive NV12 / YUY2 and convert it yourself)

take one frame from each of the two paths and compare them pixel by pixel.

How to read the difference:
  Take |A.R - B.R|, |A.G - B.G|, |A.B - B.B| for each pixel
  Report the maximum, and the percentage of pixels above a threshold

Do not expect an exact match here. There are two reasons.

  • The Source Reader’s video processing may be performing the chroma upsampling by something other than nearest neighbor. The manual implementation in 5.7. is a minimal one that uses the shared chroma directly for all 4 pixels, so differences show up most at edges
  • Rounding and intermediate precision are handled differently

So what to look at is not “does it match” but the pattern in how the differences appear.

Difference you see What to suspect
Flat areas match, differences appear only at color boundaries A difference in chroma upsampling. Expected
Everything is uniformly off The matrix (601 / 709) or the range (16..235 / 0..255) is mixed up
Banding, or a diagonal skew A hard-coded stride. See 7.2. and 7.5.
Red and blue are swapped BGRA and RGBA are mixed up
Everything looks transparent / pure black The 4th byte is not filled with 0xFF. See 7.1.

Looking at the “shape” of the difference narrows down where to look quite a lot. If the whole image is uniformly off, it is the formula or the color information; if it is localized, it is an index or the stride.

The two-stage verificationDiagram showing the two-stage verification - first pass a single known Y/U/V pixel and compare it against hand calculation, then compare the Pattern A and Pattern B frames taken from the same time in the same video and narrow down what to suspect from the shape of the difference.Stage 1: compare one pixel against hand calculationStage 2: compare frames from the two pathsNarrow down what to suspect from the shape of the differenceNeeds neither a video nor Media Foundation

Figure 19: Clear the single-pixel check first, then compare whole frames, and the shape of the difference points at the cause.

6. Which One Should You Choose?

When in doubt, the following table sorts things out quite well.

Aspect Automatic conversion (MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING) Manual conversion
Implementation speed Excellent Fair
Extracting a few stills Excellent Good
High frame volume / real-time Fair Excellent
Explicit control of matrix / range Fair Excellent
Combining with GPU / D3D Fair Good to excellent
Output formats other than RGB32 Fair Excellent
Understanding the fundamentals Good Excellent

For your first implementation, this framing makes it easy.

  • Want it working first -> automatic conversion
  • Want to own color and performance -> manual conversion

In practice, the sequence “first confirm a correct picture with automatic conversion, then replace it with the manual path” is also quite effective. If you take on everything from the start, it becomes hard to tell where the picture broke.

An order of work that pays off in practiceDiagram showing the practical order of work - confirm a correct picture with the automatic conversion first and then replace it with your own manual path, which makes it easier to tell where the picture broke.Confirm a correct picture with automatic conversion firstThen replace it with the manual pathEasier to isolate where it brokeTake on everything from the startHard to tell where it broke

Figure 20: Securing a known-good picture before moving to your own implementation makes it much easier to isolate what broke.

7. Pitfalls That Are Easy to Hit in Practice

7.1. Assuming RGB32 Is RGBA with Alpha

In memory, RGB32 is B, G, R, Alpha or Don't Care. If you write it out to a PNG as BGRA as is, the 4th byte may be 0, making the image transparent. It is safer to set it to 0xFF before saving.

7.2. Hard-Coding the Stride as width * bytesPerPixel

A very common mistake. The actual sample buffer can contain padding, so the rule is to use the actual stride to move between rows.

7.3. Confusing MF_MT_DEFAULT_STRIDE with the Actual Pitch

MF_MT_DEFAULT_STRIDE is “the minimum stride when that format is represented in contiguous memory.” For the actual pitch of the sample buffer, prefer the value returned by IMF2DBuffer::Lock2D. (pitch is another name for stride. As noted in 3.2., this article uses them with the same meaning.)

7.4. Silently Guessing 601 / 709 Without Looking at the Color Metadata

Color bugs are hard to see. They do not crash either. That is what makes them troublesome.

  • MF_MT_YUV_MATRIX
  • MF_MT_VIDEO_NOMINAL_RANGE

At the very least, look at these. And the right attitude is roughly: values your code does not support should be errors.

7.5. Locating the NV12 UV Plane with width * height

The plane offset is determined by the actual stride and height. Not by width * height. Do this sloppily and you get shifted colors or corrupted images.

How to locate the NV12 plane boundaryDiagram showing that the start of the NV12 UV plane is determined by the actual stride multiplied by the height, and that cutting at width times height leads to color drift and corrupted images.Advance stride × heightStart of the buffer (Y plane)Start of the UV planeCut at width × heightColor drift and corrupted images

Figure 21: Locate the UV plane boundary with stride x height, never with width x height.

7.6. Processing Interlaced Video Assuming Progressive

The manual samples in this article assume progressive video. Reading interlaced content as if each frame were a single field can produce comb-like artifacts. If you need deinterlacing, it is more natural to consider the Source Reader’s automatic video processing or the Video Processor MFT.

7.7. Ignoring the Quality of 4:2:0 Chroma Upsampling

For clarity, the NV12 conversion in this article uses the shared chroma directly for each pixel. That is sufficient for many uses, but if image quality is the priority, it is worth studying the upconversion approach described in the recommended YUV formats documentation.

8. Summary

When converting YUV to RGB with Media Foundation, keeping the following framework in mind makes it much harder to get lost.

  • Behind the decoder, NV12 or YUY2 — not RGB — is what normally comes out
  • If you want the easy path, request RGB32 via MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING
  • If you want control, receive NV12 / YUY2 and convert to BGRA yourself
  • On the manual path, get sampling / range / matrix / stride right before worrying about the formula
  • Being vague about BT.601 / BT.709, 16..235, and 4:2:0 / 4:2:2 leads to color drift or broken pictures

YUV -> RGB is a bit unapproachable at first. But once the picture of

  • NV12 shares U/V across 2x2 blocks
  • YUY2 shares U/V across 2 horizontal pixels
  • apply the matrix to that U/V together with Y

settles into your head, it becomes quite tame. Those mysterious cosmic-colored byte sequences start to look like properly meaningful pixels.

The picture to keep in your headDiagram showing that once the picture of NV12 sharing U/V across 2x2 blocks, YUY2 sharing U/V across 2 horizontal pixels, and applying the matrix to that U/V together with Y settles in, a mysterious byte sequence starts to look like meaningful pixels.NV12: U/V shared across 2x2Apply the matrix to that U/V and YYUY2: U/V shared across 2 horizontal pixelsThe byte sequence starts to look like meaningful pixels

Figure 22: Once you hold those two pictures - the sharing unit and the matrix - YUV byte sequences read plainly.

9. References

Sample Code for This Article

Microsoft Learn

Recent articles sharing the same tags. Deepen your understanding with closely related topics.

These topic pages place the article in a broader service and decision context.

This article connects naturally to the following service pages.

Windows App Development

This topic covers Media Foundation, the Source Reader, image saving, and video frame conversion — a Windows media-processing implementation theme that fits well with our Windows application development service.

Technical Consulting & Design Review

If you want to sort out the division of responsibility for YUV / RGB conversion, color spaces, stride, and conversion-path design up front, this topic works well as a technical consulting / design review engagement.

Frequently Asked Questions

Common questions about the topic of this article.

Why does a Media Foundation decoder output YUV instead of RGB?
Because the human eye is more sensitive to fine detail in brightness than in color, video benefits from a design that keeps Y (the brightness-oriented component) at full detail and U/V (the color-difference components) coarser. That is why, in the Windows video world, the uncompressed frames coming out of a decoder are normally YUV-family formats such as NV12 or YUY2. It also helps to read YUV as effectively meaning Y'CbCr in the context of digital video.
What is the easiest way to get RGB frames?
Enable MF_SOURCE_READER_ENABLE_VIDEO_PROCESSING on IMFSourceReader and request MFVideoFormat_RGB32. For extracting a few still images or generating thumbnails, this is by far the easiest route. However, this automatic conversion is software processing and is not optimized for real-time playback, so if you need high-volume processing or control over color, receive the frames as YUV and convert them yourself.
What should I watch out for when converting YUV to RGB myself?
It is not over once you multiply by three coefficients: subsampling (4:2:0 / 4:2:2), range, matrix, and stride all come into play. The things that most often break colors in practice are not looking at MF_MT_YUV_MATRIX and MF_MT_VIDEO_NOMINAL_RANGE, and assuming the stride is width times bytesPerPixel. The shortest path is to understand the structure of NV12 and YUY2 properly first.
What is the difference between NV12 and YUY2?
NV12 is a 4:2:0 format: a Y plane followed by a UV plane in which U and V alternate, and the 4 pixels of a 2x2 block share one U/V pair. YUY2 is a 4:2:2 format in which 2 horizontal pixels share one U/V pair. Both show up often in practice, and what differs is how much the color is thinned out (subsampling).

Author Profile

Profile page for the article author.

Go Komura

Representative of KomuraSoft LLC

Focused on Windows software development, technical consulting, and investigations into failures that are difficult to reproduce.

Back to the Blog