MPEG-TS Container Format - Audio and Video Fundamentals
作者:罗上文,微信:Loken1,公众号:FFmpeg弦外之音
First, what is MPEG? MPEG stands for Moving Picture Experts Group. This expert group is part of the Joint Technical Committee (JTC1), which was established by ISO (the International Organization for Standardization) and IEC (the International Electrotechnical Commission). JTC1 is responsible for information technology. Within JTC1, subcommittee SG29 is responsible for "audio, picture coding, and multimedia and hypermedia information". SG29 contains several working groups, including JPEG (the Joint Photographic Experts Group) and WG11, the working group responsible for moving-picture compression. Therefore, MPEG can be considered ISO/IEC JTC1/SG29/WG11, established in 1988.
It is enough to think of MPEG as an organization. MPEG mainly develops standards for audio/video compression and transmission. The MPEG organization has produced five standards so far: MPEG-1, MPEG-2, MPEG-4, MPEG-7, and MPEG-21. The MPEG-TS container format is defined in the MPEG-2 standard.
Here is a brief introduction to MPEG-2. MPEG-2 is widely used for Internet transmission protocols and, historically, for cable digital television, terrestrial digital television, satellite television, DVB, DVD, and more.
The MPEG-2 standard is divided into 10 parts, collectively known as the ISO/IEC 13818 international standard, titled "GENERIC CODING OF MOVING PICTURES AND ASSOCIATED AUDIO".
| Coding | Title | Description |
|---|---|---|
| 13818-1 | System (system) | Describes how multiple audio streams, video streams, and elementary data streams are combined into transport streams and program streams. |
| 13818-2 | Video (video) | Describes video coding methods. |
| 13818-3 | Audio (audio) | Describes audio coding methods that are backward-compatible with the MPEG-1 audio standard. |
| 13818-4 | Compliance (conformance) | Describes how to determine whether an encoded stream conforms to the MPEG-2 stream specification. |
| 13818-5 | Software (software) | Describes software implementations of parts 1, 2, and 3 of the MPEG-2 standard. |
| 13818-6 | DSM-CC (command and control) | Describes the session signaling set between servers and users in interactive multimedia networks. |
The remaining four parts of MPEG-2 are not listed because they are not widely used.
The table shows that the definition of the MPEG-TS container format is in 13818-1 (System). In 1990, the ATM video coding expert group cooperated with MPEG and turned part 13818-1 of ISO/IEC 13818 into ITU-T Rec. H.220 (System), and 13818-2 into ITU-T Rec. H.262 (Video). ITU-T Rec. H.220 and ITU-T Rec. H.262 are both part of the ITU-T standards.
These organizations cooperate with one another. They extend or adapt standards so that they fit applications in particular domains. In other words, they divide the work: each organization has its own areas of focus, but those areas overlap.
That covers the background of MPEG. Now let us discuss the MPEG-TS container format. Files whose names end in .ts use the MPEG-TS container format; ts stands for Transport Stream.
There is also an MPEG-PS container format; ps stands for Program Stream. Note that Program means program in the sense of a broadcast program, not a computer program. PS is mainly used in environments where errors are unlikely to occur, such as DVD discs.
TS packets have a fixed length of 188 bytes, while PS packets have variable lengths. The fixed packet structure of a TS stream gives it strong resistance to transmission errors.
What is a transmission error?
Errors arise during signal transmission because attenuation changes the signal voltage, damaging the signal in transit and producing bit errors. Noise, pulses caused by alternating current or lightning, transmission equipment failures, and other factors can all cause bit errors (for example, a transmitted 1 is received as 0, or vice versa). For various reasons, errors are unavoidable when digital signals are transmitted. -- Baidu Baike
In simple terms, transmission errors are a low-level problem. Digital signals are binary streams such as 111000, and a transmission error occurs when a 1 becomes a 0 in transit. This is also why the TCP and UDP protocols have a checksum field: when an error is detected, the lower layer discards the packet instead of passing it to the application layer.
In practice, for UDP or TCP, using TS or PS provides the same resistance to transmission errors, because protocols such as UDP and TCP handle errors for you. However, the TS and PS standards are not only used with UDP and TCP. Devices in the digital-video domain, such as video recorders and DVD players, also use PS and TS.
Both TS and PS were originally used mainly in digital television. China's digital television standard is DVB, which stands for Digital Video Broadcasting.
When we used to watch television, there were many channels. A receiving antenna could switch between channels and receive different broadcasts. A channel contains many programs. Pay attention to this concept of a program: the English word is Program, and this term appears frequently in MPEG-2 documents.
For example, at a given time the CCTV channel may contain programs CCTV1 through CCTV14. Each program has an audio stream and a video stream, as shown in the diagram:

How are these channels, programs, and audio/video streams distinguished inside TS? This is where the TS container format becomes complicated. Early televisions (receivers) had limited performance and could not communicate with the base station; they could only passively receive the digital signal broadcast by the base station.
The data transmitted over the Internet using TCP/UDP is a digital signal, as is the digital signal broadcast by a digital-TV base station: binary data such as 001001. However, when TS is used on the Internet, it usually carries one video stream and one audio stream, so many TS fields are unnecessary for us. Internet TS is a relatively simple use case. A TS stream containing one program is called a Single Program Transport Stream (SPTS).
Readers developing set-top boxes or similar devices and wanting to understand digital television in depth can read Video Demystified: A Handbook For The Digital Engineer.
The general background is now covered. Let us move on to practical work. As usual, we need software that can parse the TS format. This article uses Elecard Stream Analyzer, which is available as a free 30-day trial. The sample file is juren.ts, located in the source directory.
Elecard Stream Analyzer is strongly recommended; it is an excellent tool, and Elecard's software is worth trying.
Open juren.ts in Stream Analyzer. The screenshot is shown below:

Let us start with the basic types. A TS stream actually contains three types of packets.
ES, Elementary Stream: You can think of an ES packet as an H.264-encoded video frame or an audio frame, although ES is not limited to audio/video frames.
PES, Packetized Elementary Stream: An additional layer is wrapped around the ES packet, adding information such as PTS and DTS.
TS, Transport Packet: A fixed 188-byte packet. A large PES packet is split into multiple smaller pieces and transmitted inside TS packets.
We will analyze the fields of a TS packet from the outside in. See the image below:

Note that Stream Analyzer does not parse the fields in byte order. The fields in its interface are not displayed in byte order.
As shown above, a Transport Packet is a TS packet. These packets all begin with 0x47, which makes synchronization possible in some scenarios. The ISO/IEC 13818-1 standard provides syntax for parsing a TS packet, as shown below:

The sync_byte in the image is 0x47. transport_packet() is a function for parsing a TS packet. Its syntax is shown below:
Pay attention to the word syntax. Many MPEG documents provide syntax written in a C-like language. If you are familiar with C, their pseudocode is easy to understand. Syntax is pseudocode in this context.

The syntax above shows that the beginning of a TS packet contains the following fields:
- sync_byte, the synchronization field, fixed at 0x47. Position: bits 0-7. If the same value 0x47 appears elsewhere but is not a synchronization field, it would presumably need escaping; I have not examined this closely, so this is left as an open question.
Additional note: 0x47 probably does not need escaping. The TS length is fixed at 188 bytes; see the parsing code for the details.
transport_error_indicator, the transmission-error flag. In a TCP/UDP scenario, this field should always be 0.
playload_unit_start_indicator, the payload-start flag. Because one PES packet is split across multiple TS packets, a marker is needed to identify the beginning.
transport_priority, the transport priority. This field is probably not used much on the Internet and can generally be ignored.
PID. P is not an abbreviation for Program, and I do not know what it stands for. PID identifies the type of data in the payload. The standard explains this as follows:

transport_scrambling_control, a restriction field. It is generally used for paid programs. In the past, satellite television required a card and a subscription fee to watch certain programs. This field is rarely used in Internet scenarios.
adaptation_field_control, the adaptation-field flag. This field has four values. 10 and 11 indicate that additional fields follow and must be parsed; 01 and 11 indicate that this TS packet has a payload.

continuity_counter. If a TS packet contains a payload, this field increments. It does not seem to be used much in Internet scenarios, and I am not sure exactly what it does, so this remains an open question for now.
data_byte. The payload data should begin here. The syntax in the document processes the data in a loop for N iterations; subtracting the adaptation_field length from 184 gives N. You can think of N as the payload size. In a real implementation, the code does not necessarily loop exactly N times.
We can manually parse the six fields above with Notepad++, as shown below:

The first byte of a TS file is always 0x47. In an Internet scenario, this sync_byte is actually a redundant compatibility field. If we designed a new container format specifically for TCP/UDP, 0x47 would not be needed. When TCP carries an m3u8 TS live stream, it may not transmit this 0x47. I will capture packets and refine this section later.
In the ATSC standard, the 0x47 synchronization byte is never encoded or transmitted. Instead, a specific two-level synchronization pulse replaces the byte during transmission, and the receiver inserts the 0x47 synchronization byte at this position.
The next two bytes, bytes 1-2, are 0x40 0x11. These two bytes are composed of the four fields transport_error_indicator, playload_unit_start_indicator, transport_priority, and PID.
Convert 0x40 0x11 to binary: 0100 0000 0001 0001. Read from left to right; there is no bit 0 here, so counting starts at 1. The first bit is 0, meaning transport_error_indicator is 0 and there is no transmission error. The second bit is 1, the playload_unit_start_indicator, meaning that this is the first payload packet. The third bit is 0, meaning transport_priority is 0.
The remaining 13 bits are the PID value. 0 0000 0001 0001 is the PID, namely 17.

Now look at the third byte. Its value is 0x10, or 0001 0000 in binary. The third byte consists of transport_scrambling_control, adaptation_field_control, and continuity_counter, as shown below:

Finally, look at the fourth byte. Its data_byte value is 00. This completes the explanation of the eight fields in the TS packet header.
The bslbf in the Table 2-3 image stands for Bit string, left bit first. uimsbf stands for Unsigned integer, most significant bit first. Both terms are defined in the standard documents.
Next, let us inspect the remaining bytes, as shown below:

The ff bytes shown above are padding bytes used to fill the packet to 188 bytes. For Internet traffic, this is essentially waste. The TS container format was originally designed for uses other than the Internet.
How do we know that these bytes contain a Service Description Table? It is probably because the PID in the header is 0x11. The parsing syntax for the Service Description Table is shown below:

The first TS packet in juren.ts is actually custom data. Its PID is 0x11 and its table_id is 0x42, which identifies it as User private. I will state plainly what this initial TS packet is for: after looking at it for a long time, I still could not understand it. These syntax documents are best studied together with TS parsing code, such as the TS code in FFmpeg.
Now let us examine the second TS packet, as shown below:

The PID and table_id shown above are both 0, so this TS packet contains a PAT (Program Association Table). A reminder: earlier we said that TS packets encapsulate PES packets, but TS is not limited to PES. TS can also encapsulate PSI data. PSI stands for Program Specific Information, which you can think of as information specific to a program. PSI is not one table; it is a general term. PAT, PMT, CAT, and NIT are all PSI tables, as shown below:

That is a brief introduction to PSI. The TS container format is genuinely complex, with many tables, so readers should study it in depth alongside the source code.
Next, let us find all TS packets belonging to the first video frame and use them to understand the TS container format. We inspect the AVPacket in Qt Creator, as shown below:

In FFmpeg's AVPacket, the data pointer points to the actual H.264-encoded data, and the pos field points to the beginning of the TS packet. You can see the characteristic TS bytes 0x47 0x41.
The first video frame is 3768 bytes. Its beginning is 00 00 01 09 f0 00 00 00 00 06 00 07, and its ending is ab 28 fd 45 f9 30 42 22 70 34 00 00 03 00 00 03 00 02 1e 00.
Since 3768 bytes is large, the frame must be stored in multiple TS packets. Let us find all of those packets, as shown below:

There are several points to note in the image above.
I counted approximately 20 TS packets. These 20 TS packets together form one video frame.
Of the 20 TS packets, only the first has playload_unit_start_indicator equal to 1; all the others have 0.
The continuity_counter value keeps increasing.
This completes the analysis of the MPEG-TS container format. Only the basics were covered. Readers are encouraged to read the ISO/IEC 13818-1 standard, which is well written, and to study the format in depth alongside FFmpeg's TS code.
Here is a tip for reading standards: do not think that reading English documents is difficult. Since DeepL became available, reading English technical documents has become much easier. I also do not read just one document; I combine Chinese translations and Chinese analysis articles to understand a standard. Since Chinese is our native language, Chinese material is usually enough to understand a standard. If Chinese material is genuinely scarce, I read English articles, but that is much slower. Even with DeepL, some sections still require careful study.
Finally, a piece of supplementary knowledge:
- MPEG-PS and MPEG-TS are both built on PES, so the two formats can be converted into each other.