IRC log of simile on 2003-09-26
Timestamps are in UTC.
- 16:00:05 [RRSAgent]
- RRSAgent has joined #simile
- 16:00:14 [Rob]
- AndyS, I just type '/invite RRSAgent'
- 16:00:23 [Rob]
- it might be a client thing
- 16:01:01 [AndyS]
- Hmm - so did I ... in ViRC
- 16:01:29 [ericm]
- ericm has joined #simile
- 16:01:45 [ericm]
- ericm has changed the topic to: simile sept 26 telecon
- 16:01:47 [AndyS]
- I just get error messages: not enough parameters
- 16:03:02 [jse]
- I never have enough parameters
- 16:03:39 [mickBass]
- mickBass has joined #simile
- 16:03:53 [marbut]
- marbut has joined #simile
- 16:04:29 [marbut]
- mickBass: David has serious problems with calls at this time now.
- 16:04:46 [marbut]
- So I suggested he spoke to Eric, as he is the other person with a tightly constrained schedule
- 16:05:08 [marbut]
- action: dk,em - sort out time for call
- 16:05:38 [marbut]
- I want to talk about the next plenary. We've got a plenary scheduled for November.
- 16:06:07 [marbut]
- Unfortunately I have a prior commitment on November 7th. So we need to work out if we want to continue
- 16:06:28 [marbut]
- with the plenary, and we can delegate this to other people e.g. Nick, Martin, Mark
- 16:06:45 [marbut]
- The alternative is to delay it by four weeks
- 16:07:23 [marbut]
- ms: did we have specific aims for the plenary?
- 16:07:41 [marbut]
- ks: I think moving the plenary back a month probably won't help Mick
- 16:08:08 [marbut]
- ms: moving it forward in October, that's not so good. The real value of these meetings is getting beyond sticking points.
- 16:08:35 [marbut]
- At the moment I feel we are moving. And of course its not just the meeting, its the followup
- 16:08:43 [marbut]
- Is that stuff you can delegate
- 16:09:11 [marbut]
- mickBass: I'd request that Mark, Martin and Nick would get involved. Nick would take my role.
- 16:10:02 [marbut]
- MS: we are still doing the cleanup from the financial logistics from the last one, so we need to think do we really need this
- 16:10:43 [marbut]
- mickBass: My sense is we are probably going to need the time. If our goal is to get something together by late December / early January, we will need some time together.
- 16:11:18 [marbut]
- KS: We need to walkthrough the demo script and see how things are going to fit together.
- 16:11:29 [marbut]
- AndyS: are we going to have any hires on board by then?
- 16:11:58 [marbut]
- em: I think that was a goal here. I don't know if we will be in place, it will be close.
- 16:12:11 [marbut]
- I feel comfortable with the suggestion that John made to Mick
- 16:12:30 [marbut]
- (JE asked for a plan for the meeting)
- 16:12:46 [marbut]
- MS: I'll save the dates, we need to think about who will take this over, what we want to do then
- 16:13:14 [marbut]
- em: The PIs are a critical resource at this meeting. So another thing we can do is a less inclusive meeting than the plenary,
- 16:13:23 [marbut]
- perhaps involving the PIs the week of the 20th
- 16:13:34 [marbut]
- MS: We need people working on the demonstrator there
- 16:13:57 [marbut]
- The version I looked at was still pretty high level, it didn't get into details, we could do something that week
- 16:13:57 [jse]
- w/o *Oct* 20th?
- 16:14:20 [marbut]
- em: oct 20th is hard, there is an international SW conference in florida all week
- 16:14:33 [marbut]
- could do week before, week after
- 16:14:55 [marbut]
- is anybody going to that conference, apart from me? (silence) I take that as a no
- 16:15:18 [AndyS]
- http://iswc2003.semanticweb.org/
- 16:15:38 [marbut]
- mickBass: its conceivable I could make cambridge on the 15th of october
- 16:16:13 [marbut]
- ms: why don't the PIs take this offline?
- 16:17:39 [marbut]
- mickBass: next item - corpus
- 16:17:48 [marbut]
- em: no updates
- 16:18:19 [marbut]
- ms: I spoke to the Artstor folks again, I'm going to give them a SIMILE presentation next tuesday
- 16:18:47 [marbut]
- I'm hoping I can leave there with the data, perhaps that day, this should help
- 16:19:43 [marbut]
- mark: what are the possibility of getting the Getty thesauri?
- 16:19:59 [marbut]
- em: you can download a portion of it. There are conditions of use for this.
- 16:20:22 [ericm]
- http://www.getty.edu/research/conducting_research/vocabularies/license.html
- 16:20:23 [marbut]
- Mark: What about the licensing terms?
- 16:20:41 [ericm]
- http://www.getty.edu/research/conducting_research/vocabularies/download.html
- 16:21:36 [marbut]
- mickBass: Its a question of how much money - we are underspent on the project
- 16:21:47 [marbut]
- em: I can find out how much it might cost
- 16:22:45 [marbut]
- em: it should give you an understanding of how the thesaurus works at a high level
- 16:22:55 [marbut]
- but both the aat, tgn and ulan would be of interest
- 16:23:51 [marbut]
- ms: the library of congress lcsh (subject headings)
- 16:24:01 [marbut]
- ks: there is a CIA place name one
- 16:24:17 [marbut]
- em: that's already in RDF
- 16:25:36 [marbut]
- mickBass: what about wordnet?
- 16:25:46 [marbut]
- em: yes, there is an RDF version of wordnet also
- 16:25:59 [ericm]
- Wordnet in RDF : http://xmlns.com/2001/08/wordnet/
- 16:26:51 [marbut]
- jg: xml representation of sw is a barrier to its adoption
- 16:27:11 [marbut]
- also we noticed that software to simplify creation could help with the SIMILE use cases
- 16:27:46 [marbut]
- we found some reports that reviewed existing tools, but they either didn't review all the relevant tools
- 16:27:49 [marbut]
- or they were out of date
- 16:28:08 [marbut]
- so we decided to look at tools for schemas, ontologies and thesauri. The point is these are all very
- 16:28:12 [marbut]
- similar things.
- 16:28:21 [marbut]
- We had some conclusions here:
- 16:28:40 [marbut]
- firstly that working at either the XML serialisation, or the graph level are too low level for most people
- 16:28:51 [marbut]
- people want to work at the conceptual level, like protege
- 16:29:12 [marbut]
- second the terminology is a bit obscure for people who aren't knowledge engineers
- 16:29:29 [marbut]
- so we might want to hide some of the richness from naive users
- 16:29:48 [marbut]
- third software used tabs to break down the task. This seemed a good approach
- 16:30:16 [marbut]
- fourth we need to draw a distinction between element sets & controlled vocabularies - in rdf type languages they are both classes
- 16:30:24 [marbut]
- but librarians think of them differently
- 16:30:50 [marbut]
- fifth existing tools don't provide help for schema modelling
- 16:31:09 [marbut]
- (david karger joins)
- 16:32:12 [marbut]
- mickBass: so the point from four is that its helpful to distinguish between classes - they can be collections of properties or terms in a vocabulary
- 16:32:41 [marbut]
- johng: so to carry one, there are issues here about data modelling and normalisation, and how to do it
- 16:33:10 [marbut]
- six, we only found one tool that supported repositories and supported reuse of ontologies and schemas
- 16:33:19 [marbut]
- this is the MEG project by UKOLN
- 16:33:44 [marbut]
- seven, tree representations were very common, but people did use different approaches
- 16:34:02 [marbut]
- eight, tools need to provide visualisations on the underlying data
- 16:34:21 [marbut]
- nine, the tools don't yet support multiple typing, which is a bit limiting
- 16:34:39 [marbut]
- AndyS: isn't there an interaction here between trees and conceptual modelling
- 16:35:01 [marbut]
- once you have a bunch of records, when you add the relations, its no longer tree like
- 16:35:30 [marbut]
- jg: yes, the tools have to convert lattice representations of sub / super class hiearchies into trees
- 16:35:40 [marbut]
- (similiar to the approach used in rdftwig)
- 16:36:01 [marbut]
- the next slide tries to propose a workflow for dealing with heterogeneous schemas and metadata
- 16:36:24 [marbut]
- and that we are dealing with people who may or may not have data format
- 16:36:38 [marbut]
- the diagram summarises the lifecycle presented in a previous report
- 16:37:35 [marbut]
- so the conclusions from this:
- 16:37:53 [marbut]
- take something like protege, add a faceted search rather than just a tree index, and also provide
- 16:38:14 [marbut]
- search on the free text using a tool like lucene. so we get to search metadata and free text
- 16:38:48 [marbut]
- paul: one of the ideas raised at a recent meeting was giving feedback on how people are using schemas
- 16:39:12 [marbut]
- jg: this kind of system would help. It's a first step, getting them in the same place
- 16:39:33 [marbut]
- we do some facets, other things are links as the relationships can't be displayed by facets
- 16:40:07 [marbut]
- then based on this I have put together a small web based application that does this with some existing schemas
- 16:40:25 [marbut]
- em: so you have a prototype of this?
- 16:41:24 [marbut]
- jg: its not a plug-in for protege, its a java servlet
- 16:41:37 [marbut]
- em: the point paul made about use metrics, statistics etc.
- 16:42:04 [marbut]
- when working on dc registries a while ago, use statistics, indications of policy and persistant on vocabularies
- 16:42:22 [marbut]
- were all useful metrics that help the next community decide what to invest in
- 16:42:46 [marbut]
- so under the umbrella of vocabulary reuse / discovery, adding some information to help the user decide what to choose
- 16:42:49 [marbut]
- would help
- 16:43:15 [marbut]
- I think if we could provide a system here, I think it would be a useful semantic web bootstrapping tool
- 16:43:26 [marbut]
- mickBass: other questions? comments?
- 16:45:34 [marbut]
- em: another quick comment - first thanks. second you've done some analysis, so I wonder
- 16:46:27 [jse]
- URL for SHAME editor: http://sourceforge.net/projects/shame/
- 16:46:48 [marbut]
- mickBass: any other conclusions
- 16:47:18 [marbut]
- em: I'm not sure if we are updating the website?
- 16:47:27 [marbut]
- action: update website
- 16:47:51 [marbut]
- action update website mark butler
- 16:48:08 [AndyS]
- SHAME Home page is: http://kmr.nada.kth.se/shame/
- 16:48:32 [marbut]
- mickBass: the last item, we are going to talk about demo script
- 16:49:36 [marbut]
- ms: theres a disconnect between the examples and what I thought would be in the demo
- 16:51:48 [marbut]
- ms: we can look at this as a short term image, or look into the future
- 16:52:43 [marbut]
- so in the ocw, individual items do have IMS metadata, its just they are not exposing it?
- 16:53:00 [marbut]
- but I hear what you are saying - people may package images in certain ways?
- 16:53:10 [marbut]
- but that will make the demo much more complicated.
- 16:53:21 [marbut]
- mickBass: even if we focus only on lom's that are images
- 16:53:33 [marbut]
- ms: I think we should only focus on lom's that are images. period
- 16:54:08 [marbut]
- mickBass: so in one way we simplify complexity, because the repository is just images
- 16:54:49 [marbut]
- ms: i think just dealing with mapping vocabularies is hard enough, without having to map between totally different things
- 16:55:12 [marbut]
- does the demo have to reflect how the world is now, or does it have to be in future?
- 16:55:44 [marbut]
- ms: the way I invisaged it is you have a large collection of images in vra, the other images with ims data
- 16:55:56 [marbut]
- so its a simple scenario, buts its hard enough to do that
- 16:56:23 [marbut]
- mickBass: but even in a corpus that only contains images, can we extract the information just using schema synonyms?
- 16:56:54 [marbut]
- or do we have something more complicated, where we need to understand the relationships between entities?
- 16:57:25 [marbut]
- ms: in IMS, there is a lot of inheritance that goes on, so we can decide for our search engine whether we want to do that inferencing or not
- 16:58:13 [marbut]
- you'd assume either the course name, or the subject, or type is mentioned somewhere
- 16:58:45 [marbut]
- mickBass: can you count on it being a specific location, or will it be distributed in the record. So you draw from a number of fields
- 16:58:59 [marbut]
- (that was mackenzie smith)
- 17:00:07 [marbut]
- mark: is there anyway we can get some samples of ims, so we can better understand how this would work>
- 17:00:40 [marbut]
- ms: yes, I could ask them, but at the moment they are locked up in the content management system
- 17:01:04 [marbut]
- we can probably make a couple of records
- 17:01:35 [marbut]
- mickBass: so one thread is to talk to the OCW folks, another is to try to rework the example records that Mark has created
- 17:02:11 [marbut]
- ms: I think Mark's records were too real world, we have more latitude
- 17:02:26 [marbut]
- em: I'd like to see Mark and MacKenzie work on this, get some examples together
- 17:02:52 [marbut]
- because then we can go back to OCW, and say if you provide the data in this way, this is what we can do
- 17:03:14 [marbut]
- ms: people apply vra consistently, but IMS is a lot harder, people apply it differently
- 17:03:41 [marbut]
- em: the problem with IMS its a bit cart before the horse, they have data but don't know what to do with it
- 17:03:51 [marbut]
- ms: yes, but we need to remember we are researching SIMILE not IMS
- 17:04:03 [marbut]
- mickBass: bearing in mind the time, I suggest we break there
- 17:04:54 [marbut]
- how do I save the telecon minutes?
- 17:05:30 [ericm]
- rrsagent, pointer?
- 17:05:30 [RRSAgent]
- See http://www.w3.org/2003/09/26-simile-irc#T17-05-30
- 17:06:17 [ericm]
- ok, irs logs are now world readable
- 18:00:32 [mick]
- mick has joined #simile