IRC log of simile on 2003-09-26

Timestamps are in UTC.

16:00:05 [RRSAgent]
RRSAgent has joined #simile
16:00:14 [Rob]
AndyS, I just type '/invite RRSAgent'
16:00:23 [Rob]
it might be a client thing
16:01:01 [AndyS]
Hmm - so did I ... in ViRC
16:01:29 [ericm]
ericm has joined #simile
16:01:45 [ericm]
ericm has changed the topic to: simile sept 26 telecon
16:01:47 [AndyS]
I just get error messages: not enough parameters
16:03:02 [jse]
I never have enough parameters
16:03:39 [mickBass]
mickBass has joined #simile
16:03:53 [marbut]
marbut has joined #simile
16:04:29 [marbut]
mickBass: David has serious problems with calls at this time now.
16:04:46 [marbut]
So I suggested he spoke to Eric, as he is the other person with a tightly constrained schedule
16:05:08 [marbut]
action: dk,em - sort out time for call
16:05:38 [marbut]
I want to talk about the next plenary. We've got a plenary scheduled for November.
16:06:07 [marbut]
Unfortunately I have a prior commitment on November 7th. So we need to work out if we want to continue
16:06:28 [marbut]
with the plenary, and we can delegate this to other people e.g. Nick, Martin, Mark
16:06:45 [marbut]
The alternative is to delay it by four weeks
16:07:23 [marbut]
ms: did we have specific aims for the plenary?
16:07:41 [marbut]
ks: I think moving the plenary back a month probably won't help Mick
16:08:08 [marbut]
ms: moving it forward in October, that's not so good. The real value of these meetings is getting beyond sticking points.
16:08:35 [marbut]
At the moment I feel we are moving. And of course its not just the meeting, its the followup
16:08:43 [marbut]
Is that stuff you can delegate
16:09:11 [marbut]
mickBass: I'd request that Mark, Martin and Nick would get involved. Nick would take my role.
16:10:02 [marbut]
MS: we are still doing the cleanup from the financial logistics from the last one, so we need to think do we really need this
16:10:43 [marbut]
mickBass: My sense is we are probably going to need the time. If our goal is to get something together by late December / early January, we will need some time together.
16:11:18 [marbut]
KS: We need to walkthrough the demo script and see how things are going to fit together.
16:11:29 [marbut]
AndyS: are we going to have any hires on board by then?
16:11:58 [marbut]
em: I think that was a goal here. I don't know if we will be in place, it will be close.
16:12:11 [marbut]
I feel comfortable with the suggestion that John made to Mick
16:12:30 [marbut]
(JE asked for a plan for the meeting)
16:12:46 [marbut]
MS: I'll save the dates, we need to think about who will take this over, what we want to do then
16:13:14 [marbut]
em: The PIs are a critical resource at this meeting. So another thing we can do is a less inclusive meeting than the plenary,
16:13:23 [marbut]
perhaps involving the PIs the week of the 20th
16:13:34 [marbut]
MS: We need people working on the demonstrator there
16:13:57 [marbut]
The version I looked at was still pretty high level, it didn't get into details, we could do something that week
16:13:57 [jse]
w/o *Oct* 20th?
16:14:20 [marbut]
em: oct 20th is hard, there is an international SW conference in florida all week
16:14:33 [marbut]
could do week before, week after
16:14:55 [marbut]
is anybody going to that conference, apart from me? (silence) I take that as a no
16:15:18 [AndyS]
http://iswc2003.semanticweb.org/
16:15:38 [marbut]
mickBass: its conceivable I could make cambridge on the 15th of october
16:16:13 [marbut]
ms: why don't the PIs take this offline?
16:17:39 [marbut]
mickBass: next item - corpus
16:17:48 [marbut]
em: no updates
16:18:19 [marbut]
ms: I spoke to the Artstor folks again, I'm going to give them a SIMILE presentation next tuesday
16:18:47 [marbut]
I'm hoping I can leave there with the data, perhaps that day, this should help
16:19:43 [marbut]
mark: what are the possibility of getting the Getty thesauri?
16:19:59 [marbut]
em: you can download a portion of it. There are conditions of use for this.
16:20:22 [ericm]
http://www.getty.edu/research/conducting_research/vocabularies/license.html
16:20:23 [marbut]
Mark: What about the licensing terms?
16:20:41 [ericm]
http://www.getty.edu/research/conducting_research/vocabularies/download.html
16:21:36 [marbut]
mickBass: Its a question of how much money - we are underspent on the project
16:21:47 [marbut]
em: I can find out how much it might cost
16:22:45 [marbut]
em: it should give you an understanding of how the thesaurus works at a high level
16:22:55 [marbut]
but both the aat, tgn and ulan would be of interest
16:23:51 [marbut]
ms: the library of congress lcsh (subject headings)
16:24:01 [marbut]
ks: there is a CIA place name one
16:24:17 [marbut]
em: that's already in RDF
16:25:36 [marbut]
mickBass: what about wordnet?
16:25:46 [marbut]
em: yes, there is an RDF version of wordnet also
16:25:59 [ericm]
Wordnet in RDF : http://xmlns.com/2001/08/wordnet/
16:26:51 [marbut]
jg: xml representation of sw is a barrier to its adoption
16:27:11 [marbut]
also we noticed that software to simplify creation could help with the SIMILE use cases
16:27:46 [marbut]
we found some reports that reviewed existing tools, but they either didn't review all the relevant tools
16:27:49 [marbut]
or they were out of date
16:28:08 [marbut]
so we decided to look at tools for schemas, ontologies and thesauri. The point is these are all very
16:28:12 [marbut]
similar things.
16:28:21 [marbut]
We had some conclusions here:
16:28:40 [marbut]
firstly that working at either the XML serialisation, or the graph level are too low level for most people
16:28:51 [marbut]
people want to work at the conceptual level, like protege
16:29:12 [marbut]
second the terminology is a bit obscure for people who aren't knowledge engineers
16:29:29 [marbut]
so we might want to hide some of the richness from naive users
16:29:48 [marbut]
third software used tabs to break down the task. This seemed a good approach
16:30:16 [marbut]
fourth we need to draw a distinction between element sets & controlled vocabularies - in rdf type languages they are both classes
16:30:24 [marbut]
but librarians think of them differently
16:30:50 [marbut]
fifth existing tools don't provide help for schema modelling
16:31:09 [marbut]
(david karger joins)
16:32:12 [marbut]
mickBass: so the point from four is that its helpful to distinguish between classes - they can be collections of properties or terms in a vocabulary
16:32:41 [marbut]
johng: so to carry one, there are issues here about data modelling and normalisation, and how to do it
16:33:10 [marbut]
six, we only found one tool that supported repositories and supported reuse of ontologies and schemas
16:33:19 [marbut]
this is the MEG project by UKOLN
16:33:44 [marbut]
seven, tree representations were very common, but people did use different approaches
16:34:02 [marbut]
eight, tools need to provide visualisations on the underlying data
16:34:21 [marbut]
nine, the tools don't yet support multiple typing, which is a bit limiting
16:34:39 [marbut]
AndyS: isn't there an interaction here between trees and conceptual modelling
16:35:01 [marbut]
once you have a bunch of records, when you add the relations, its no longer tree like
16:35:30 [marbut]
jg: yes, the tools have to convert lattice representations of sub / super class hiearchies into trees
16:35:40 [marbut]
(similiar to the approach used in rdftwig)
16:36:01 [marbut]
the next slide tries to propose a workflow for dealing with heterogeneous schemas and metadata
16:36:24 [marbut]
and that we are dealing with people who may or may not have data format
16:36:38 [marbut]
the diagram summarises the lifecycle presented in a previous report
16:37:35 [marbut]
so the conclusions from this:
16:37:53 [marbut]
take something like protege, add a faceted search rather than just a tree index, and also provide
16:38:14 [marbut]
search on the free text using a tool like lucene. so we get to search metadata and free text
16:38:48 [marbut]
paul: one of the ideas raised at a recent meeting was giving feedback on how people are using schemas
16:39:12 [marbut]
jg: this kind of system would help. It's a first step, getting them in the same place
16:39:33 [marbut]
we do some facets, other things are links as the relationships can't be displayed by facets
16:40:07 [marbut]
then based on this I have put together a small web based application that does this with some existing schemas
16:40:25 [marbut]
em: so you have a prototype of this?
16:41:24 [marbut]
jg: its not a plug-in for protege, its a java servlet
16:41:37 [marbut]
em: the point paul made about use metrics, statistics etc.
16:42:04 [marbut]
when working on dc registries a while ago, use statistics, indications of policy and persistant on vocabularies
16:42:22 [marbut]
were all useful metrics that help the next community decide what to invest in
16:42:46 [marbut]
so under the umbrella of vocabulary reuse / discovery, adding some information to help the user decide what to choose
16:42:49 [marbut]
would help
16:43:15 [marbut]
I think if we could provide a system here, I think it would be a useful semantic web bootstrapping tool
16:43:26 [marbut]
mickBass: other questions? comments?
16:45:34 [marbut]
em: another quick comment - first thanks. second you've done some analysis, so I wonder
16:46:27 [jse]
URL for SHAME editor: http://sourceforge.net/projects/shame/
16:46:48 [marbut]
mickBass: any other conclusions
16:47:18 [marbut]
em: I'm not sure if we are updating the website?
16:47:27 [marbut]
action: update website
16:47:51 [marbut]
action update website mark butler
16:48:08 [AndyS]
SHAME Home page is: http://kmr.nada.kth.se/shame/
16:48:32 [marbut]
mickBass: the last item, we are going to talk about demo script
16:49:36 [marbut]
ms: theres a disconnect between the examples and what I thought would be in the demo
16:51:48 [marbut]
ms: we can look at this as a short term image, or look into the future
16:52:43 [marbut]
so in the ocw, individual items do have IMS metadata, its just they are not exposing it?
16:53:00 [marbut]
but I hear what you are saying - people may package images in certain ways?
16:53:10 [marbut]
but that will make the demo much more complicated.
16:53:21 [marbut]
mickBass: even if we focus only on lom's that are images
16:53:33 [marbut]
ms: I think we should only focus on lom's that are images. period
16:54:08 [marbut]
mickBass: so in one way we simplify complexity, because the repository is just images
16:54:49 [marbut]
ms: i think just dealing with mapping vocabularies is hard enough, without having to map between totally different things
16:55:12 [marbut]
does the demo have to reflect how the world is now, or does it have to be in future?
16:55:44 [marbut]
ms: the way I invisaged it is you have a large collection of images in vra, the other images with ims data
16:55:56 [marbut]
so its a simple scenario, buts its hard enough to do that
16:56:23 [marbut]
mickBass: but even in a corpus that only contains images, can we extract the information just using schema synonyms?
16:56:54 [marbut]
or do we have something more complicated, where we need to understand the relationships between entities?
16:57:25 [marbut]
ms: in IMS, there is a lot of inheritance that goes on, so we can decide for our search engine whether we want to do that inferencing or not
16:58:13 [marbut]
you'd assume either the course name, or the subject, or type is mentioned somewhere
16:58:45 [marbut]
mickBass: can you count on it being a specific location, or will it be distributed in the record. So you draw from a number of fields
16:58:59 [marbut]
(that was mackenzie smith)
17:00:07 [marbut]
mark: is there anyway we can get some samples of ims, so we can better understand how this would work>
17:00:40 [marbut]
ms: yes, I could ask them, but at the moment they are locked up in the content management system
17:01:04 [marbut]
we can probably make a couple of records
17:01:35 [marbut]
mickBass: so one thread is to talk to the OCW folks, another is to try to rework the example records that Mark has created
17:02:11 [marbut]
ms: I think Mark's records were too real world, we have more latitude
17:02:26 [marbut]
em: I'd like to see Mark and MacKenzie work on this, get some examples together
17:02:52 [marbut]
because then we can go back to OCW, and say if you provide the data in this way, this is what we can do
17:03:14 [marbut]
ms: people apply vra consistently, but IMS is a lot harder, people apply it differently
17:03:41 [marbut]
em: the problem with IMS its a bit cart before the horse, they have data but don't know what to do with it
17:03:51 [marbut]
ms: yes, but we need to remember we are researching SIMILE not IMS
17:04:03 [marbut]
mickBass: bearing in mind the time, I suggest we break there
17:04:54 [marbut]
how do I save the telecon minutes?
17:05:30 [ericm]
rrsagent, pointer?
17:05:30 [RRSAgent]
See http://www.w3.org/2003/09/26-simile-irc#T17-05-30
17:06:17 [ericm]
ok, irs logs are now world readable
18:00:32 [mick]
mick has joined #simile