The Anatomy of an Undeclared Caption Track

A video stream ships one caption track and the player offers two. The extra one is real, decodable, and nowhere in the manifest. Tracing it took me into the bytes hiding inside H.264 frames, into a four line gap in hls.js, and eventually into writing a CEA‑608 encoder to build a stream that could fail on demand. Budget ~15 minutes.
1. A menu with a language nobody shipped
hls.js#4920 is four years old, two sentences long, and labelled good first issue. A stream declares one closed caption rendition, English on CC1. The player presents two, the second one Spanish. The reporter attached a test stream and moved on.
It reads like a cosmetic complaint. It isn’t, and the reason is worth stating before any code: a caption menu is a contract, not an inventory. A publisher lists the caption tracks they have prepared, reviewed, and are willing to put in front of a viewer. A player that lists something else is not being generous, it is breaking the contract on the publisher’s behalf.
What actually shows up in that extra slot, in the wild, is rarely a polished second language. It is a stale channel from an old encoder pass, or a technician’s scratchpad, or bytes that decode into garbage. The player labels it Spanish or Unknown CC because those are its defaults, and now an accessibility menu contains an entry nobody wrote, nobody reviewed, and nobody can turn off.
2. Where 608 captions actually live
The first surprise for anyone arriving from WebVTT: CEA‑608 captions are not a file. There is no sidecar to fetch, no separate playlist, no URL. The caption bytes ride inside the video frames themselves, in an H.264 SEI message (the ATSC A/53 flavour), two bytes per frame, decoded by the player as it decodes picture.
Two bytes per frame is a punishingly small pipe, and the format spends it well. Those two bytes carry either a control code or a character pair, and each field of the signal multiplexes two independent channels. Field one carries CC1 and CC2. A control code with bit 3 set switches the decoder to channel two; everything after it belongs to CC2 until something switches it back.
Which means a single stream of bytes, with no framing and no table of contents, can hold English and Spanish at once. Step through a real one:
What you just saw: one byte stream, no container, no metadata, filling two separate caption buffers because a control code said so. Note what is not in there: any statement about which of these channels a viewer is supposed to be offered. The wire format has no opinion. It carries whatever the encoder put on it, including whatever an encoder put on it three years ago and forgot.
3. What the manifest promises
That opinion lives one level up, in the Multivariant Playlist. HLS has a tag for exactly this, and its whole job is to name the channels the publisher intends to present:
One rendition. INSTREAM-ID="CC1" is the publisher saying: of the channels the video may physically contain, this is the one that is yours. The media can carry four; this playlist offers one.
So the correct behaviour is not ambiguous, and it does not require guessing intent. The manifest declares, the media carries, and the manifest wins. Everything below is about a player that reads the declaration and then does not use it.
4. Labels, not gates
hls.js keeps four fixed slots for 608, one per channel, built once when its TimelineController is constructed. Each slot starts with a default label and language code from config:
Those defaults are the reason the phantom track in the bug report is named Spanish. Nothing in that stream is Spanish. Slot two is simply called Spanish before anyone asks.
Then two things happen, and they never speak to each other.
The playlist arrives. onManifestLoaded walks the CLOSED-CAPTIONS renditions, matches each INSTREAM-ID to a slot, and fills in the real label, the language, and the rendition object itself:
Read that last line again. slot.media is a perfect record of this channel was declared. The information the fix needs is already in the object, already populated, already correct.
Caption bytes decode. Separately, whenever the 608 parser produces a cue on any channel, OutputFilter calls one method:
And createCaptionsTrack creates the track. It does not look at slot.media. It does not look at the playlist at all. Channel produced bytes, therefore track exists.
That is the entire bug, and it is not a missing feature: the declaration was parsed, stored, and used only for naming. The library had the answer in hand and consulted it for the label while ignoring it for the decision. Which is also why the issue sat for four years looking cosmetic. Someone reading the code sees INSTREAM-ID handled right there in onManifestLoaded and reasonably assumes it is wired up.
5. Making the declaration load bearing
The fix is a predicate, a derived flag, and two guards. The flag records whether the playlist declared anything, because that distinguishes two cases that must behave differently: a playlist that declared CC1 and not CC2, versus a playlist that declared no captions at all and cannot be used to filter anything.
That flag is a getter rather than a stored field, and it did not start out that way. Section 9 covers why the maintainer asked for the change, because the reason generalises well past this file.
The obvious guard goes at the top of createCaptionsTrack. The second one is the interesting one, and it is not optional. addCues runs for every cue, and in native rendering mode it does this:
cueCache[trackName] is only ever populated by track creation. Block creation and leave this path alone, and the first undeclared cue dereferences undefined and throws inside caption parsing. One decision, enforced at both places a channel can become visible.
Then a detail that is easy to miss until a second stream loads: captionsProperties lives as long as the player instance, but media belongs to one playlist. Call loadSource() again and yesterday’s declarations are still sitting in those slots, filtering today’s stream. So onManifestLoading clears them.
Set the manifest and the build below and watch what the player ends up offering:
What you just saw: the change bites in exactly one square of that matrix, the one where a playlist declares some channels and the media carries more. Declare both and both survive. Declare none and the old behaviour stays, because a playlist that declares nothing is not evidence that undeclared channels are unwanted, it is the absence of evidence. Same for a Media Playlist loaded directly, which has no way to declare renditions at all.
That third row is why the change ships behind filterUndeclaredClosedCaptions, defaulting to true. Filtering is the spec reading and the reported expectation, but somebody out there is quietly relying on an undeclared channel today, and a one line escape hatch is cheaper than an argument.
6. The stream that had to be built
Here is where a caption bug stops being like other bugs. I had a fix and a unit test. What I did not have was a way to watch it fail, because the test stream in the issue had gone offline sometime in the four years since it was filed. No stream, no reproduction, no proof. And a fix nobody has seen fail is a fix nobody has seen work.
The requirement is specific: video whose manifest declares CC1 only, and whose frames carry both CC1 and CC2. Nothing public was going to hand me that, so I built it.
Step one is the part I expected to be easy. ffmpeg does not have a CEA‑608 encoder. It decodes 608 happily, and its -a53cc flag forwards caption side data that already exists on decoded frames, but there is no path from a subtitle file to 608 bytes in an H.264 SEI. So the caption bytes had to be generated by hand: odd parity in the top bit of each 7 bit code, preamble address codes for row and style, pop‑on sequences (resume caption loading, erase non‑displayed memory, write, end of caption), and the channel two variants of every control code with bit 3 set.
Every pair in the first lab on this page came out of that encoder. It is the reference implementation for the thing being tested, which is a slightly uncomfortable place to stand, so it gets checked against a decoder that was not written by me: ffmpeg reads the finished stream back and prints the captions.
Step two: get the bytes into the video. Two bytes per frame, in an SEI NAL, ahead of the first slice of each access unit. Encode with access unit delimiters so frame boundaries are trivial to find, walk the Annex B stream, and splice:
Emulation prevention is the part that bites if you skip it: any 00 00 0x sequence in the payload has to have a 0x03 spliced in, or a decoder reading the stream finds a start code that was never meant to be there. Then remux to MPEG‑TS, segment, and write three master playlists that declare different subsets of the same media.
Thirty seconds of video, three hundred caption pairs, and the build asserts that every one of them landed. The generator lives in the tester repo linked at the bottom, and it is reusable for anything else that needs a 608 stream with specific channel content, which is a thing I now know is annoyingly hard to find.
7. Proof you can look at
Unit tests prove the predicate. They cannot prove that a browser, a demuxer, a 608 parser, and a text track API all end up somewhere different because of it. So the last artifact is a page that runs both libraries at once: hls.js at the pinned commit on the left, the same commit plus the patch on the right, same manifest, same segments, same wall clock.
Building it that way was deliberate. Two builds from one commit means the only variable on the page is the diff. If the panes differ, the patch is the reason; there is no other candidate.
Left pane: two caption tracks, and selecting the second one paints Spanish text over video whose manifest declares English and nothing else. Right pane: one track. The scenario switcher covers the guardrails too, because the interesting claim is not only that the fix filters, it is that the fix filters and nothing else changes.
That is the part I care about as a tester. The failing case is easy to demonstrate once you have a stream. The four cases that must keep working are the ones that decide whether a patch is safe to merge.
8. Shipping it upstream
What went to video-dev/hls.js: thirty lines in timeline-controller.ts, two in config.ts, five unit tests, an API.md entry for the new option, and the regenerated api‑extractor report. Two of the five tests fail on master, which is the only way to know a test is testing anything.
The rest of the work, the encoder, the injector, the split screen page, is not in the pull request and should not be. It links from the description instead. A maintainer reviewing 125 lines of library change should not have to read an ffmpeg pipeline to accept it, but they should be able to click through to one if they want to know how the evidence was made.
9. What the review changed
The first review came back changes requested, opening with “Makes sense. Mostly nit-pick change requests.” Five comments, every one with the exact diff attached. Nothing disputed what the patch did. All five were about its shape.
Four were genuine nits: rename a method, update its two call sites, avoid a for...in loop, extract a repeated object literal into a helper. Worth doing, not worth writing about.
The fifth one was not a nit, and it is the reason this section exists. I had written the declaration flag as a field:
Set it to true when a rendition is parsed, back to false on playlist reset. Two assignment sites, one boolean, perfectly readable. His note: “Please remove this flag.”
The argument is that the fact was already in the object. A track having .media set is the declaration, so a separate boolean recording “something was declared” is a second copy of information the class already held. And two copies of one fact can disagree.
They disagree quietly, too. Any future code path that populates captionsProperties without remembering to set the flag, or clears one without the other, breaks the filter in the direction where nothing throws: captions stop being filtered and the phantom track comes back. No error, no test failure unless someone wrote exactly the right test, just the original bug wearing a different hat.
A getter cannot drift, because there is only one place the answer can come from. That is the whole idea, and it cost four lines.
The part I did not expect was the compiler doing the demolition. Delete the field, add the getter, and TypeScript immediately flags both assignment sites: Cannot assign to ‘captionsDeclared’ because it is a read-only property. It pointed straight at every line that had been maintaining the duplicate, which were exactly the lines that could have drifted. Then the reset collapsed too, because replacing captionsProperties with a fresh default object clears the declarations and the derived flag in one assignment.
Net effect of the review round: 16 fewer lines, one less thing that can be wrong, and no behaviour change at all. The rename inverted a boolean and rearranged a condition, so I checked all four input combinations by hand before agreeing to it; both forms reduce to the same expression, and all 1189 unit tests passed unchanged afterwards.
Two of the five edits could have compiled clean while behaving wrong: dropping the ! from a call site would have inverted the filter, and deleting the flag reset without putting the new one in place would have leaked declarations across playlists. Both typecheck perfectly. Neither is the kind of thing a compiler has an opinion about.
10. Check yourself
1 · A stream carries CEA‑608 on CC1 and CC2 and its playlist declares neither. How many caption tracks should a conforming player present?
2 · The player labels an undeclared channel “Spanish”. Where did that name come from?
captionsTextTrack2Label, a config default applied to slot two before any manifest is parsed. Nothing in the stream claims to be Spanish. Channels three and four get Unknown CC the same way.