Introduction to Container Formats
作者:罗上文,微信:Loken1,公众号:FFmpeg弦外之音
Before discussing container formats, let us first explain what a file format is. Computers store everything as binary data, including ASCII strings; one ASCII character occupies one byte. An int stored in a file is also binary data, and one int occupies 4 bytes.
When a commonly used text editor opens an ASCII file, it displays normally. But if it opens a binary file that stores int values, the result looks strange because the text editor interprets the 4 bytes of each int as four ASCII characters.
It is not only int; many other data types can also be stored in files. Therefore, here I divide file formats into two types:
String storage: open the file directly with an ordinary text editor such as Notepad++ or Sublime to view its contents.
Binary storage: specialized software is needed to analyze the file contents.
In fact, audio and video container formats are a type of binary file format. Their special characteristic is reflected in the English term mux (multiplexing). Anyone who has studied TCP has probably heard of multiplexing. Audio and video containers work the same way: they can combine multiple streams into one file. This is a container format; this is muxing.
The encoding and compression systems output individual streams: the video coding system outputs a video stream, and the audio coding system outputs an audio stream. A container format combines multiple streams into one file.
FLV, MP4, and MKV are all container formats, and the contents stored in these files are binary. In practice, a container format is a standard defined by organizations or companies: which fields a file in this format contains, what each field means, how many bytes it occupies, where it is located in the file, and how the fields are nested. These are the things a container format defines.
A container format is essentially a standard. For example, as long as everyone reads and writes data according to the MP4 standard document, there will be no problem. Imagine agreeing with someone that bread goes in the first drawer and medicine in the second. The other person knows to look in the first drawer for the bread. If I do not put the bread in the first drawer as agreed, the other person cannot find it and their parser reports an error.
The main purpose of a container format is to provide an agreed set of rules. A container format can combine audio and video streams.
There are many different container formats. Different formats are defined to solve problems in different scenarios. For example:
Video-on-demand: the video has already been recorded and placed on a server, and the client fetches a small portion on demand for playback instead of downloading the entire video. MP4 is well suited to this scenario because it defines an
sttsindex table and related data structures, allowing fast seeking. For example, seeking to a particular timestamp is much faster in MP4 than in FLV. (Note: FLV can add akeyframeindexto speed up seeking.)Live streaming: the box structure of an MP4 file cannot be generated until the entire video has been recorded, while a live stream has no known end time. FLV is a progressive format and is therefore well suited to live streaming.
Thus, different container formats solve problems in different scenarios. Containers also solve a problem common to all scenarios: audio-video synchronization. A container assigns a PTS timestamp to every video and audio frame.
When a player plays only a video stream or only an audio stream, it does not need these PTS timestamps: the video stream can be played at its frame rate and the audio stream at its sample rate. PTS helps synchronize audio and video.
Here is why. Most operating systems we use are time-sharing systems, meaning that each task is allocated a certain CPU time slice.
Suppose Windows is playing a video without an audio stream at 24 frames per second, and the current time is 8:00:00 p.m. The first video frame starts at 8:00:00:00. At this frame rate, the second frame must be displayed 41 ms later, at 8:00:00:41, and the third at 8:00:00:82 for smooth playback. However, because this is a time-sharing system, the CPU may be busy with another task at 8:00:00:41 and unable to switch back to the player thread. The second frame might not start until 8:00:00:61, 20 ms late. What can we do in this situation? Even if we developed the player, there is nothing we can do when the CPU is temporarily overloaded. It is enough for the third frame to be 20 ms late as well. If the CPU is idle afterward, playback will still look smooth. This is the single-video-stream case.
Now suppose an audio stream and a video stream are played simultaneously. The picture and sound should be synchronized; for example, when the second video frame plays, the second audio frame should play too. But if the video stream is 20 ms late and the audio playback thread ignores the video thread's stall, the audio stream continues at its own sample rate and plays 20 ms ahead of the video. This difference keeps accumulating, and the audio and video eventually become badly out of sync. Therefore, the container must assign a PTS to every audio and video frame. Each playback thread observes the PTS values already played by the other stream to decide whether to continue, sleep, or drop a frame.