Linked Data for Language Technology (LD4LT) Group Kick-Off Meeting and Roadmap meeting, 21 March, Athens, Greece
Linked Data (LD) has proven beneficial in many new and unforeseen ways for Language Technology (LT) and the newly gained interoperability and availability of LT data and services is currently receiving industry adoption. With the foundation of the LD4LT W3C community group, we would like to start the discussion and analyse current trends as well as offer a crystallization point to coordinate the development of future LD-based LT applications. See the agenda.
All feedback is welcome and participation is open to all interested organisations and individuals from industry and academia.
The LD4LT Group Kick-Off and Roadmap Meeting is supported by the LIDER project, the MultilingualWeb community, the NLP2RDF project, the Working Group for Open Data in Linguistics as well as the DBpedia Project.
As input to the discussion and the work of the LD4LT group, you may consider to fill in the first LIDER survey. During the kick-off meeting, via the survey and in the LD4LT group, provide your view on how linked data and language technology should benefit each other.
To be held 7-8 May 2014 in Madrid, Spain, W3C announced today the seventh MultilingualWeb workshop in a series of events exploring the mechanisms and processes needed to ensure that the World Wide Web lives up to its potential around the world and across barriers of language and culture.
This workshop is made possible by the generous support of the LIDER project. As part of the event, LIDER will organize a roadmapping workshop on linked data and content analytics.
Anyone may attend all sessions at no charge and the W3C welcomes participation by both speakers and non-speaking attendees. Early registration is encouraged due to limited space.
Building on the success of six highly regarded previous workshops, this workshop will emphasize new technology developments that lead to new opportunities for the Multilingual Web. The workshop brings together participants interested in the best practices and standards needed to help content creators, localizers, language tools developers, and others meet the challenges of the multilingual Web. It provides further opportunities for networking across communities. We are particularly interested in speakers who can demonstrate novel solutions for reaching out to a global, multilingual audience.
ITS 2.0 provides a foundation for integrating automated processing of human language into core Web technologies. ITS 2.0 bears many commonalities with its predecessor, ITS 1.0, but provides additional concepts that are designed to foster the automated creation and processing of multilingual Web content.
Work on application scenarios for ITS 2.0 and gathering of usage and implementation experience will now take place in the ITS Interest Group.
The Unicode Consortium has announced Version 6.3 of the Unicode Standard and with it, significantly improved bidirectional behavior. The updated Version 6.3 Unicode Bidirectional Algorithm now ensures that pairs of parentheses and brackets have consistent layout and provides a mechanism for isolating runs of text.
Based on contributions from major browser developers, the updated Bidirectional Algorithm and five new bidi format characters will improve the display of text for hundreds of millions of users of Arabic, Hebrew, Persian, Urdu, and many others. The display and positioning of parentheses will better match the normal behavior that users expect. By using the new methods for isolating runs of text, software will be able to construct messages from different sources without jumbling the order of characters. The new bidi format characters correspond to features in markup (such as in CSS). Overall, these improvements also bring greater interoperability and an improved ability for inserting text and assembling user interface elements.
The improvements come with new rigor: the Consortium now offers two reference implementations and greatly improved testing and test data.
In a major enhancement for CJK usage, this new version adds standardized variation sequences for all 1,002 CJK compatibility ideographs. These sequences address a well-known issue of the CJK compatibility ideographs — that they could change their appearance when any process normalized the text. Using the new standardized variation sequences allows authors to write text which will preserve the specific required shapes of these CJK ideographs, even under Unicode normalization.
Version 6.3 includes other improvements as well:
- Improved Unihan data to better align with ISO/IEC 10646
- Better support for Hebrew word break behavior and for ideographic space in line breaking
The MultilingualWeb-LT Working Group has published a Proposed Recommendation of Internationalization Tag Set (ITS) Version 2.0. The technology described in this document provides a foundation for to integrating automated processing of human language into core Web technologies. ITS 2.0 bears many commonalities with its predecessor, ITS 1.0 but provides additional concepts that are designed to foster the automated creation and processing of multilingual Web content. ITS 2.0 focuses on HTML, XML-based formats in general, and can leverage processing based on the XML Localization Interchange File Format (XLIFF), as well as the Natural Language Processing Interchange Format (NIF). Comments are welcome through 22 October.
The MultilingualWeb-LT Working Group has published a Last Call Working Draft of Internationalization Tag Set (ITS) Version 2.0. ITS 2.0 makes it easier to integrate automated processing of human language into core Web technologies. ITS 2.0 focuses on HTML, XML-based formats in general, and can leverage processing based on the XML Localization Interchange File Format (XLIFF), as well as the Natural Language Processing Interchange Format (NIF). Comments are welcome through 10 September.
The two-day workshop surveyed and shared information about currently available best practices and standards that can help content creators and localizers address the needs of the multilingual Web, including the Semantic Web. Attendees also heard about gaps that need to be addressed, and enjoyed opportunities to network and share information between the various different communities involved in enabling the multilingual Web. Half of the second day was dedicated to an Open Space discussion with breakouts.
The workshop was sponsored by the EU-funded QTLaunchPad project and Verisign. “During the W3C Rome Workshop we were able to identify the gaps holding organizations back from achieving a truly multilingual Internet and discuss best practices that will further our collective goal of achieving a web void of cultural borders”, said Pat Kane, senior vice president of Naming and Directory Services at Verisign.
The recently announced Internationalization Tag Set 2.0 showcase event in Dublin now allows for remote participation. Please register by 17 June 6 p.m. UTC. We will provide dial in details to registered
participants. The number of remote participants is limited and we choose on a
first-come, first-served basis – get your seat soon!
On 18 June the MultilingualWeb-LT Working Group holds a showcase event in Dublin about the upcoming Internationalization Tag Set (ITS) 2.0 specification. Group participants demonstrate implementations for authoring ITS 2.0 data categories, for using them in localization workflows, and for improving machine translation or other language technology processes with ITS 2.0. Participation is free, but registration is required.
The draft implements all changes since the previous publication of 11 April 2013. There are no remaining open issues. The Working Group is planning to finalize ITS 2.0 now: this is your last time to provide feedback! The Last Call period ends 11 June.
ITS 2.0 provides metadata to foster the adoption of the multilingual Web.