13:57:53 RRSAgent has joined #me 13:57:57 logging to https://www.w3.org/2026/07/07-me-irc 13:57:57 Zakim has joined #me 14:00:28 ohmata has joined #me 14:00:44 Meeting: MEIG Monthly Meeting 14:00:49 Chair: Chris, Song 14:01:01 nigel has joined #me 14:01:43 Present: Chris Needham, Song Xu, Kazuyuki Ashimura, Nigel Megitt, Yuta Hagio, Hisayuki Ohmata, Nishitha Dey, Wolfgang Schildbach, Rob Smith 14:01:48 Chair+ Wolfgang 14:02:27 Present+ Dimitri Poborksi 14:03:29 wschildbach has joined #me 14:04:17 RobSmith has joined #me 14:05:36 present+ Louay Bassbouss 14:06:15 chrisn: We will continue the discussion about AI and Media Entertainment today. 14:06:37 hagio_nhk has joined #me 14:06:38 Topic: AI media use cases 14:06:38 ... next TPAC meeting is in Dublin, will be looking for agenda topics (later) 14:06:48 Louay has joined #me 14:07:00 present+ Louay_Bassbouss 14:07:32 ... Mr. Zheng and Song presented AVS standardization use cases, like 2D 3D image converison, volumetric use cases etc. 14:07:32 rrsagent, make log public 14:07:36 scribe+ nigel 14:07:37 rrsagent, draft minutes 14:07:38 I have made the request to generate https://www.w3.org/2026/07/07-me-minutes.html kaz 14:08:10 Song: AVS in China use cases were shared, and a demo in limited time. 14:08:52 ... we could break down specific features / characteristics and categorize in different levels. Then describe impact on MEIG. 14:09:35 present+ Atsushi Shimono 14:09:47 chrisn: Sounds good, and we should gather input from more members here on what they are interested in (AI usecases as they apply to media) 14:09:57 present- Louay_Bassbouss 14:10:23 rrsagent, draft minutes 14:10:24 I have made the request to generate https://www.w3.org/2026/07/07-me-minutes.html kaz 14:10:30 ... so we have as broad a set of usecases as possible. We are interested to capture your usecases. 14:10:56 Louay: MOQ could be a topic for the TPAC meeting as well. 14:11:29 hagio has joined #me 14:11:39 ... is related to AI for the stream, if you have AI chatbots or voice assistants, or avatars, that are generated in the backend and rendered in the browser 14:11:54 ... that would make MOQ relevant as well in the AI context. 14:12:32 q+ (can be later) to mention my action on new requirements for subtitles enabled by AI 14:12:32 cpn: Let's capture the AVS categorization. 14:12:42 q+ to mention my action on new requirements for subtitles enabled by AI (can be later) 14:13:26 Song: when categorizing use cases, I am thinking end to end use cases. Codecs are relevant, and transportation. 14:14:04 ... believe that what Roy mentioned is also relevant for discussion. Would be appreciated if transportation could be included. 14:14:48 ... github has the original information from dom, the intention is to see if there are changes in priority 14:15:29 ... With Chris' response about labelling, I started; my understanding of the domain is that there are several groups of AI efforts. 14:16:14 ... c2pa is first area; relevant for content authored by humans and assisted by AI; or by AI alone 14:17:09 ... these impacts can be group 1: text based content generation. 14:17:38 ... beyond that, for MEIG, more complicated impacts on webcodecs, a/v, webvr etc., this can be second group. 14:18:39 ... three use cases: next-gen codecs driven by neural networks; enhanced experience with immersive experience (volumetric, super resolution, ...) 14:19:01 cpn: do you know who drives MPEG DCVC? 14:19:36 q? 14:19:52 Song: Tech people from Microsoft research centers AFAIK; within MPEG, I observed that some of attendees talk about DCVC but have not started official process. 14:20:26 ... third sub-usecase is multimodal agents combined with text and A/V on the web. 14:20:43 ... if we add content labelling scenario, that would be four use cases. 14:20:56 cpn: We may have more to add... 14:21:14 present+ Chris Seeger 14:21:49 Song: in general, I believe the use cases 1 and 2 have similar foundations, in that they are upgrades of existing codecs to/with AI, or enhancements of the user experience 14:22:21 ... They have features in common, for example raw data for rendering is same/similar. 14:22:55 ... webGPU and webNN enhance the codecs 14:23:29 ... volumetric could lead another direction: higher data rates, higher demands on bandwidth, lower latency streaming for web transport 14:24:46 ... for use cases 3, they are the most emerging use cases; people use ChatGPT, Claude, Doubao, Qwen through web platform. In the future, these agents will be upgrade from text based input to multimodal 14:25:39 ... the simultaneous inference from vision and audio could demand higher requirements 14:26:18 ... realtime requires webtransport for low latency 14:26:45 ... Is there a chance for optimization or a need to integrate AI labelling? 14:27:17 ... key output transforms web media from codec to rendering to agentic interaction 14:27:40 ... webGPU and webNN/webTransport are foundational 14:27:54 cpn: what are thoughts about MOQ and WebTransport? 14:28:40 Louay: MOQ/WebTransport is the basic layer; WebRTC can be used as can be WebTransport 14:29:17 q+ 14:29:24 ... I can provide quick intro into MOQ and relation to WebRTC and the browser, summarize use cases (media distribution, data, AI voice assistance, production) 14:30:11 ... at least webTransport is required by MOQ which is built on QUIC. WebTransport is also important for realtime but can also be used with MSE but loses some control 14:30:41 scribe+ cpn 14:30:53 nigel: had an action to put together a crash course on ???, still have that action 14:31:18 cpn: we should add that to the use cases on this slide. 14:32:03 ... nigels use case is about using AI for captions and descriptive metadata like motion or loudness levels that can enrich the description of what is happening in A/V 14:33:04 Song: responding to previous topic about WebTransport, we had this before in other working groups. When we talk about WebTransport we talk about a better user experience 14:33:38 ... it could be important foundation for multimodal agents. Even though the topic is similar to earlier, it is a research item in AI 14:33:44 cpn: Agree, this is foundational. 14:34:13 ... at this stage, is it enough that we have identified it? Louay, a future look at it would be welcome, let's follow up. 14:34:16 q? 14:34:18 ack n 14:34:18 nigel, you wanted to mention my action on new requirements for subtitles enabled by AI (can be later) 14:34:49 -> https://www.w3.org/2025/10/smartagents-workshop/report.html#interoperability Smart Voice Agents Workshop Report - 25-25 Feb. 2026 14:34:54 kaz: Thank you for the information provided, these are important use cases. In February there is a workshop. 14:35:49 ... one of the topics was relevant for voice technologists. Would make sense to have another workshop on media & AI including various stake holders across the world. 14:36:26 ... target area is getting broader, makes sense to have more stake holders including browser vendors, and AI vendors. Thoughts? 14:36:32 ack k 14:37:45 Song: I got involved in the voice agent workshop; you are correct that for voice agents we organized this workshop. Could be useful to come up with a similar analysis or report, setting different targets or milestones. The use case characterization will be useful. 14:38:48 kaz: Main topic was voice agents; next time will be broader like AI bots and agents in general. Our target is even broader. a smart agent workshop would make sense. 14:39:22 song: Will review the output for the voice agent workshop and check whether results can be used for further use cases. 14:39:31 q? 14:39:56 kaz: my main point is thinking about smart agents. 14:40:22 ... and we already have three use cases for the workshop. 14:40:35 cpn: this is something that we could plan to do 14:40:48 s/smart agents/our own workshop about media and AI/ 14:41:23 RobSmith: What ties into use case 4 (content labeling), looing at this from a provacy/security angle, if you think about a still image on the web, 14:42:18 ... we see things in it but which are difficult for machines to extract. But this is becoming easier. Now, if I can recognize something then so can the AI. 14:43:00 s/Main topic was voice agents/The main topic for the Voice Agents workshop was voice agents/ 14:43:17 ... if all this data becomes available, then agents can go and label images with who is in them; or the audio. This becomes this dystopian scenario where one can search for me and the engine comes up with all locations I've ever been 14:43:53 ... this is something we should guard against. There are already robots files that guard some if this; we should think about something similar for AI. 14:44:15 ... my premise is that if I as a human can do it so can AI (but at a larger scale). 14:44:33 ... so we need to think about what data can be aggregated from all of this? 14:44:35 q? 14:44:45 ... this comes with the content labelling mentioned earlier 14:45:32 cpn: the labelling as intended here was meant as signal that content was AI generated; regulations are coming up in the EU and US that mandate this 14:45:34 s/; next time will be broader like AI bots and agents in general. Our/while our own interest is "Media and AI" including AI bots and AI agents in general. So our/ 14:45:39 rrsagent, draft minutes 14:45:41 I have made the request to generate https://www.w3.org/2026/07/07-me-minutes.html kaz 14:46:30 ... there is a spectrum from AI assisted to completed generated. The technology generates signatures to make the content tamper evident -- any editing would be visible and can be made traceable. 14:46:46 i|We will continue the discussion about AI|scribenick: wschildbach| 14:46:50 rrsagent, draft minutes 14:46:51 I have made the request to generate https://www.w3.org/2026/07/07-me-minutes.html kaz 14:47:02 ... you end up with a content provenance log. The big driver is AI. 14:47:25 ... the labeling that you are talking about is machine learning, to understand media content and generate metadata. 14:48:17 RobSmith: Yes, this is something that humans can do and so can AI. There could be an issue with, say, 1000 pictures of a person, and one can mine for commonalities 14:48:28 ... how are we protecting against this? 14:49:22 cpn: within my org, we have a large archive of media. A usecase for us would be to find all clips of a politician. And it is exactly as you describe. Same with voice, find who speaks at any one time. 14:49:40 ... what I like about your description is it is across the entire web. 14:50:33 ... at the moment, AI is trained mostly on text, but in the future there are concerns of privacy and protecting poeple's interests. 14:52:08 RobSmith: If you regard the internet as a library, finding something by text is possible using search engines. Once this is available for images/audio, more use cases become available through aggregation: who met whom? 14:52:30 ... Ai can discover these hidden connections. And every member of the public can make these inferences. 14:53:00 cpn: there is a usecase around what content should be accessible to the AI? 14:53:10 ... this is worth capturing. 14:53:24 ... and there is active work on that. 14:54:05 RobSmith: The use case could be called "Rapid Access to media"; WebVMT ties into this 14:54:31 ... an example of this: OGSC testbed was just finished. 14:55:48 ... the dashcam in my car has an accelerometer; could look for a spike which would indicate a collision. Horizontal spike: collision; vertical spike: pothole. Tie this to location and you find potholes and speed bumps 14:56:06 ... people can look at this demo. 14:56:44 cpn: This is an example of interesting annotation, worth capturing as another use case. Song, can you share slides? 14:56:47 OGC Testbed-21 Dash Cam Analysis Video: https://youtu.be/DJN-JjekmR8 14:56:55 ... does anyone else have interesting use cases? 14:58:11 Louay: capturing motion is another example of multimodal analysis. LLM systems and other components of TTS producing audio lose emotion. Maybe LLMs can be asked to provide emotion as part of the output? 14:58:56 ... putting emotion into the voice agent is important for many use cases. You need emotion in audio, and in the avatar facial expression. This is another use case. 14:59:21 cpn: we have some work to do. Song, can you capture these use cases? 14:59:39 Song: I'll go through the notes and complete the use cases. 14:59:51 cpn: we can then build on this. 15:00:26 I summarised my use case in a comment in issue#123: https://github.com/w3c/media-and-entertainment/issues/123#issuecomment-4853703464 15:01:00 ... one more thing. I updated the gh issue with our meeting times at TPAC. Please think about what topics we want to cover in the times that we have. 15:01:15 ... in September we need to pull these together and form a plan. 15:01:21 https://github.com/w3c/media-and-entertainment/issues/121 MEIG TPAC agenda 15:01:44 s/... one more thing/ cpn: one more thing/ 15:02:23 ... next meeting August 4th. In the meantime, please find the gh issues (1,2,3) and add your use cases. 15:02:42 s/(1,2,3)/123/ 15:33:15 cpn has joined #me 15:33:22 rrsagent, draft minutes 15:33:23 I have made the request to generate https://www.w3.org/2026/07/07-me-minutes.html cpn 15:33:33 rrsagent, make log public 16:50:44 Zakim has left #me 18:08:32 cabanier has joined #me