AI-Sourced Future Requirements for Timed Text Task Force

Timed text formats like TTML and WebVTT support a variety of presentation syntax and semantics, as well as some metadata entities, in some cases in an extensible way. These match current practices for authoring and displaying timed text in use cases such as captions and subtitles, amongst others, where the timed text resources relate to some external media such as audio or video.

As device capabilities have increased so has the scope for more advanced presentation approaches, including complex animations and styling. Advances in software used to process and analyse media files, including “AI” or large language model approaches, are enabling the cost-effective capture of more data. For example, where previously it was feasible to capture sound effects, speech and speaker information, it may now be possible to automate the identification of prosody, intensity and emotion, and to use novel presentation mechanisms to convey those additional dimensions to users.

The AI-Sourced Future Requirements for Timed Text Task Force (AIRFOT) will discover any new requirements that are not yet well supported by timed text formats, and assess whether any of those require changes to format specifications.

Homepage
Homepage
Task Force of
Media and Entertainment Interest Group
Shortname
airfot

Participation

Participation in this task force requires invitation by the group chairs.

Leadership

Facilitators
  • Nigel Megitt
  • Chris Needham
  • Wolfgang Schildbach
  • Song XU