12:06:53 RRSAgent has joined #pmwg 12:06:57 logging to https://www.w3.org/2026/09/10-pmwg-irc 12:06:57 RRSAgent, make logs Public 12:06:58 Meeting: Publishing Maintenance Working Group 12:07:43 ivan has changed the topic to: Meeting Details 2026-09-10: https://lists.w3.org/Archives/Public/public-pm-wg/2026Sep/0001.html 12:07:44 Chair: wendy, susan 12:07:44 Meeting: Publishing Maintenance Working Group Telco 12:07:44 Agenda: https://www.w3.org/mid/410CF6A4-3B67-468C-B984-64B941787C2A@neustudio.com 12:58:39 DaleRogers has joined #pmwg 12:58:44 GeorgeK has joined #pmwg 12:59:01 present+ 12:59:03 AvneeshSingh has joined #pmwg 12:59:06 present+ 12:59:28 sueneu has joined #pmwg 12:59:36 gautierchomel has joined #pmwg 12:59:47 toshiakikoike has joined #pmwg 12:59:56 MasakazuKitahara has joined #pmwg 12:59:58 present+ 13:00:01 wendyreid has joined #pmwg 13:00:03 present+ 13:00:13 Brady has joined #pmwg 13:00:20 present+ 13:00:22 present+ 13:00:26 kimberg has joined #pmwg 13:00:27 shiestyle has joined #pmwg 13:00:28 present+ 13:00:38 present+ 13:00:46 present+ 13:00:49 duga has joined #pmwg 13:00:53 present+ 13:01:23 present+ avneesh 13:02:39 present+ laurent 13:03:05 scribe: sueneu 13:03:08 Topic: Annotations 13:03:54 gpellegrino has joined #pmwg 13:04:04 present+ 13:04:09 CharlesL has joined #pmwg 13:04:16 present+ 13:04:22 present+ hadrien 13:04:24 Laurent: This is where we left off for the summer, but there are some issues we can tackle… 13:04:31 present+ 13:04:43 Hadrien has joined #pmwg 13:04:51 present+ 13:04:57 ...I'm still looking for implementers, Thorium will get annotations in 3.6 in 2 months as a test 13:05:12 s/I'm/... I'm/ 13:05:23 …I hope someone from Readwise or [?] will do also 13:05:41 subtopic: https://github.com/w3c/epub-specs/issues/2852 13:05:48 s/[?]/colibrio/ 13:07:13 ... This issue was offering breadcrumbs to the reading system to get back to the origin 13:07:25 q? 13:07:43 +1 13:07:53 LaurentLM has joined #pmwg 13:07:59 present+ 13:08:07 …since no one expressed a need for this I propose we remove issue 2852 13:08:39 q+ 13:08:45 ack Hadrien 13:09:26 Hadren: we have the body, and this would block other contextual elements, there wouldn't be any human readable code left if we drop this 13:09:58 Laurent: If you drop the HTML you would still have something you could export from the locator, like the publisher and the book 13:10:35 Hadrien: I am not a fan of having something specific, a general text can be most useful. It could be a title, something from the TOC 13:10:39 q+ 13:10:44 q+ 13:11:01 …what I'm describing is a little different than what you have here. You could resolve it is a few ways 13:11:54 …when you have access to the file it is possible to extract the information, but when you have only the extraction, it could be come a problem. We could use something broader in scope 13:11:58 ack ivan 13:12:44 ivan: are we talking about the same thing? On the annotation set there is an about object, and the issue isn't about this. The meta object was created on the target which means something undefined 13:13:00 Hadrien: I was talking about an annotation that is not on the annotation set level 13:13:21 …basically I'd like to make sure annotations are useful even when I don't have access to the original publication 13:13:59 ack duga 13:14:03 Laurent: If we want that it means replacing the structure with a string giving context, like chapter title or page. It would be up to the reading system to decide what that is 13:14:50 duga: I agree with Hadrien, but is this something I need to do for version 1 of the spec. It feels useful, but it will take some work to figure out the right way to do it. 13:14:56 q+ 13:15:27 ack Hadrien 13:15:28 q+ 13:15:30 …if we're not specific, this could be too obscure to be used. If we don't want to do this for 1.0 we should remove it 13:15:52 s/same thing./same thing? 13:16:14 Hadrien: we work with the publication timeline, like at all times I have a header or a progress bar, so I know where I am in a publication. Also useful in a search 13:17:01 …a publication timeline is easy to generate with prepaginated media, and audio or video. Its more complex with a reflowable ebook but it is useful. And helpful for the user. 13:17:17 s/same thing./same thing?/ 13:17:28 …this string would be similar to what we would generate for a timeline. Each reading system does it their own way 13:17:32 ack ivan 13:18:40 ivan: to be practical. even though we haven't talked about a timeline, I don't think this specification will make it to recommendation by the end of this working group (February '26). Se we can extend the timeline and come up with a proper set of attributes. 13:19:08 …we should acknowledge that this is open and needs more work. Someone could take on finding more specific properties. 13:19:17 Laurent: I am OK with this proposal 13:19:36 subtopic: https://github.com/w3c/epub-specs/issues/2884 13:20:38 Laurent: The target may not be a specific segment in the document unless we consider that the whole document can be a specific segment 13:21:34 …the problem comes from the word "segment" and what can be the size of the target. We know if can be a video or audio. The term segment comes from the W3C model itself. 13:22:05 …I proposed a change of working to "segment of interest" I'm looking for ideas to rewrite that 13:22:07 q? 13:22:19 q+ 13:22:22 s/Laurent/LaurentLM/ 13:22:56 Hadrien: by segment do you mean fragment? 13:23:45 LaurentLM: perhaps we should remove the sentence, we should just say what it is. There are restraints in our annotation beyond what is in the W3C annotations 13:23:45 ack duga 13:24:19 Duga: There are nuances here, and we can delete this sentence. We just have to explain how things are different. 13:24:32 subtopic: https://github.com/w3c/epub-specs/issues/3009 13:25:24 Ivan: This is a general thing, what should be listed as metadata for an annotation set. What we have now is not much 13:26:02 LaurentLM: you prefer we remove "generator" it was part of the W3C model, I see it as a useful tool 13:26:07 q+ 13:26:50 …you also propose to remove the DC format, because we shouldn't speak about format in this specification 13:27:33 …the properties retained are identifier, format, title, publisher, creator, date of release. These are common properties to find in epub 13:27:54 …you propose to remove format, why? 13:27:59 ack ivan 13:28:31 q+ 13:28:36 ivan: my question about generator, is to define what it should be, I don't know how I would debug this and what I would see. 13:28:42 ack duga 13:28:47 LaurentLM: I see, I agree 13:29:15 q+ 13:29:17 Duga: I am against putting debugging information in any export format, it doesn't belong in the spec it is a privacy issue 13:29:39 Ivan: that is in line with what I'm saying, put in a URL. 13:29:57 ack gautierchomel 13:29:59 Duga: If I use a screen reader, I may not export that information to everyone 13:30:22 Ivan: about the format issue, it is a question we should answer ourselves 13:31:11 q+ 13:31:35 …I have no idea how PDF works, if all the things we define are usable in PDF and we are not able to confirm this. PDF is irrelevant unless we go through the whole document and add PDF specifications 13:32:16 …that's why format isn't useful, because the only value would be EPUB 13:32:27 ack Hadrien 13:32:50 gautierchomel: I might need to know the format of the output 13:33:07 q+ 13:33:18 q+ 13:33:49 Hadrien: I don't think we can trust all the metadata that we have, we probably need less. I'm OK with title and time stamp. DC date seems weird, we may not have this information in a reliable place. 13:33:49 ack wendyreid 13:34:43 s/the format of the output/the format of the output for use case 5,3 Annotations used in the publishing workflow/ 13:35:02 wendyreid: I agree with Hadrien, there are contexts inwhich ID is important, these may be context specific, this information may be less reliable outside of that rs 13:35:28 ack ivan 13:35:48 …I see what you are saying gautierchomel, I think some of what you mention could be inferred from the file itself, and other information might escape 13:36:08 q+ 13:36:40 ivan: we are talking about all this in isolation, when I export information, these metadata values will be in the exported package, we need to say that these must be a copy of the values use in the document 13:36:49 …and this is testable for conformance 13:37:17 …the question is whether the importing system can trust it, and maybe not, but it provides information 13:37:53 LaurentLM: I agree there is an advantage to exporting information about the document in the annotation set 13:38:26 ack Hadrien 13:38:27 q+ 13:38:40 …when I am importing annotations from one system into the same title on another system, the information would be helpful. Especially if a human can make the decsion 13:39:53 q+ 13:39:56 Hadrien: we know unique identifiers are not always unique. I think this will fail. I expect the real use case will be that people will import annotations for a particular book. I don't expect annotations to translate well from one system to another. 13:39:59 ack duga 13:40:02 q+ 13:40:09 q+ 13:41:25 Duga: I agree with Hadrien about the identifier, and also that there might be information useful to the user, we may want to be able to give the user information about the annotations before they import them. 13:41:29 ack ivan 13:41:39 …some of the information is useful, but DC identifier is useless 13:41:54 q+ 13:42:26 ack wendyreid 13:42:27 ivan: what if I take an identifier like a hash of the whole document, then I can check if the origin of the annotations and the current document are the same 13:42:40 Duga: I don't know how well that will work in practice 13:43:15 q- 13:43:19 s/whole document/whole package document/ 13:43:50 wendyreid: I agree knowing dc identifier is unreliable, the uu id can be too specific, but that can be updated per edition, what if we identify the publication through multiple parts of the metadata in a decsending order. 13:44:07 s/uu id/uuid/ 13:44:52 ack Hadrien 13:44:53 …ultimately we have something that is useful to show the reader before they import it. We can use title, publisher, and more information from the reading system like related titles. 13:46:08 Hadrien: We've seen examples of people trying to use hash and it fails. One example is KO reader sync, they send a hash and a path or two paths. Whenever you do anything with the file the hash changes. 13:46:13 q+ 13:47:11 ack GeorgeK 13:47:19 …hash is a brittle thing to use, like a house of cards. I like Dugas suggestion of using the minimum information 13:47:49 GeorgeK: are we going to have enough information to create a bibliographic reference? 13:48:07 …there is more information needed than the dc:title 13:48:11 q+ 13:48:41 ack Hadrien 13:48:43 wendyreid: we are limited to what's in the packaged document, not all publishers put in good data 13:48:51 s/packaged/package/ 13:49:23 Hadrien: if we are talking about citation references, there are many styles, and it is unlikely that we would get all that data in an epub file. 13:49:54 q+ 13:49:56 LaurentLM: if we take the baseline, then we include only title and creator in the metadata 13:50:08 wendyreid: maybe publisher 13:50:10 ack DaleRogers 13:50:18 q+ 13:50:38 q+ 13:50:48 LaurentLM: publisher doesn't seem useful to the user. Even if there are two books with the same title the creator will help differentiate 13:51:26 +1 wendyreid 13:52:35 DaleRogers: from an annotation point of view, knowing which book information to put into an annotation set, does a reading system have to check the incoming information against an existing title. Is the issue at the software level or the user's level? 13:53:39 q+ 13:53:40 ack duga 13:53:46 LaurentLM: a good use case: I am a user and I want to associate annotations with a certain book. I select the epub, select the annotations, then the reading system checks the title, and if there is a discrepancy, it alerts the reader. A human choice aided by the machine 13:54:44 q+ 13:55:34 ack ivan 13:55:39 Duga: I don't know enough about dc:publisher in ebooks if it is useful or not. As a reader I don't care who published my book. But it might be important to other readers. Since this isn't usually exposed to the user, and if the data isn't good, it could be confusing. I lean toward not including it but could be talked out of it 13:56:03 q- 13:57:13 ivan: I am worried about the way we go into this discussion. We want to be sure none of the information can be misinterpreted. But then it doesn't get used because we are worried. In a large number of cases, if I know author, publication, year, etc. It will give me enough information to find it. We are throwing away everything because there are some cases where it doesn't work. 13:57:46 q+ 13:58:01 ack gautierchomel 13:58:11 …I think we should make it clear to publishers that the package document information should be there, even if in practice these things can go wrong. Let's not throw away the information because sometimes it can go wrong. 13:58:45 gautierchomel: we should encourage good practices in publishing and not punishing the ones who do it right 13:58:54 ack DaleRogers 13:59:13 LaurentLM: I made a mistake in the dc:date, it should be the date the Ebook is published 13:59:45 q+ 14:00:14 ack GeorgeK 14:00:18 DaleRogers: for an ebook we would tell our students where to get the book. I know everyone would get the same epub, same date, etc. So outputing the annotations in a classroom situation would work well. But could get more complicated in other use cases. 14:01:16 GeorgeK: students will many times not like the school reading system environment and get the book from bookshare or another place and use the reading system that works for them. Happily if the publisher provided the title to bookshare the metadata should be the same. 14:01:16 Topic: AOB 14:01:21 https://docs.google.com/document/d/1FfUjiK8PrKfqeVCAnjpgZ_rpq7ZSAziloyLT5AkZKjQ/edit?tab=t.0#heading=h.v52yw3mb2h7 14:01:46 wendyreid: we will talk about this more, here is a link to the TPAC agenda for you to review. 14:02:01 rrsagent, draft minutes