09:23:08 RRSAgent has joined #dpvcg 09:23:08 logging to https://www.w3.org/2018/12/03-dpvcg-irc 09:38:19 Ah sorry, forgot to set it up. One moment... 09:38:42 We start with an introduction round 09:38:58 Javier has joined #dpvcg 09:39:10 chair: axel polleres 09:39:27 scribe: javier 09:39:40 Niklas has joined #dpvcg 09:40:11 Eva__ULD_ has joined #dpvcg 09:40:16 Fajar has joined #dpvcg 09:40:36 topic: Welcome & Introduction 09:40:39 Martin_Kurze has joined #dpvcg 09:40:56 Webex … we wait for Bert to establish a room for us 09:42:17 Start with a quick introductory round 09:42:40 Eva from ULD Germany, working in SPECIAL project 09:42:50 Rigo, working for W3C and ERCIM 09:43:14 Bart, also from ULD, just joined the group 09:43:31 Niklas, from WU, working in the Expedite project 09:43:50 Fajar, from TU Vienna, working also in the Expedite project 09:43:54 Simon, PhD @ WU Vienna & working as research scientist at Siemens 09:44:07 Javier, from WU, working in the SPECIAL project 09:44:20 Martin, from DT, working in the SPECIAL project 09:44:27 Present: Eva Shlehan, Rigo Wenning, Bud Bruegger, Niklas Kirchner, Fajar Ekaputra, Javier Fernandez, Martin Kurze, Amr Azzam, Harshvardhan Pandit, Axel Polleres 09:44:37 Fajar, phD student in TU, working in WU in the SPECIAL project 09:45:03 s/Fajar/Amr 09:45:19 Axel, working in WU, SPECIAL project 09:45:38 Axel: the main goal is to fill the vocabularies with content 09:45:59 ... I would suggest to start from the use cases 09:46:51 Amr has joined #dpvcg 09:46:54 Axel going trough the agenda: https://www.w3.org/community/dpvcg/wiki/F2FVienna3Dec:Agenda 09:48:44 Eva: the purpose discussion is probably connected to the legal ground (tomorrow)... we can maybe start on purposes but I expect to already talk of legal ground 09:49:05 Axel: we keep this on mind, it makes sense 09:49:41 ... I agree that these concepts are all connected, so might want to connect them in the discussion 09:50:33 ... At 17 we have the meetup, one speaker is sick but maybe Rigo can jump in 09:52:14 ... as for the lunch, please send the receipt to ERCIM who will kindly sponsor it 09:53:38 Eva: Maybe we can start with an overview on where we are, also to give an update to my new colleague 09:54:12 Axel: We started the group at the same time of the start of the GDPR coming into effect 09:56:42 .... Idea was to standardise certain vocabularies around terminology used in GDPR, to describe the purposes, the different types of personal data, consent, legitime interest... 09:56:54 all starting from a personal F2F meeting here in Vienna 09:57:29 Since then we had several telcos (bi-weekly), also a meeting in the mydata conference (e.g. Elmar) 09:57:55 harsh has joined #dpvcg 09:58:47 there is an IRC channel (this), a mailing list (with an archive as well), a wiki, an internal version accessing with W3C account, the minutes of our meetings.. 09:59:10 we collect actions in the "Tracker" 09:59:37 Bud has joined #dpvcg 10:03:32 We agreed on the scope of the paper: focus on categories of personal data (we also talked of instance data, e.g dates of birth, for data portability... still to clarify), purposes for personal data hadling, processing involved, data subjects, controllers, processors and recipients, storage and security aspects (location, duration, security measures) 10:03:40 s/paper/group 10:04:08 and last, means of legitimation for personal data processing (content, legitime interest, etc.) 10:04:41 See these points at: https://lists.w3.org/Archives/Public/public-dpvcg/2018Nov/0000.html 10:05:49 Then we started to collect use cases: https://www.w3.org/community/dpvcg/wiki/Use-Cases,_Requirements,_Vocabularies 10:06:19 most of them from SPECIAL and DECODE 10:06:24 Ramisa has joined #dpvcg 10:06:37 Niklas_ has joined #dpvcg 10:07:23 We also collected different vocabularies including such terms, as we don't want to repeat efforts (and mistakes) 10:08:06 Hash: Maybe we go to the taxonomies first 10:08:32 rigo has joined #dpvcg 10:08:43 Hash: we agreed on some taxonomies: approved means approved during the telcos 10:09:23 ... I tried to identify categories and terms in each of them 10:10:06 ... There is commonality in most of them, but the structure is different 10:10:27 a very structure new ontology might not align well with this 10:10:42 I opened my personal room on webex https://mit.webex.com/meet/rigo 10:10:51 Taxonomy: https://www.w3.org/community/dpvcg/wiki/Taxonomy 10:12:37 Rigo: if you want an ontology of what personal data is, you have to look into vcard and P3P with a full range of personal data (by microsoft), related to profiling.... 10:13:09 ... but this is boling the ocean (thousands of terms). That's why we use Linked Data to start some island and connect them 10:13:30 meeting number is 640 414 996 10:13:48 Harsh: Agree, it is very challenging to create such categories 10:14:29 Niklas: agree, sometimes is more the association with the id rather than the data itself 10:14:59 Eva: is a moving target, the decision is always context dependent and depends on the link to other information 10:15:25 Niklas: it depends on the processing more than the context 10:15:51 Eva: What is approved in the wiki? 10:16:07 Axel: It is "approved for discussion", or "in scope" 10:17:00 Bud: how can one access the wiki? 10:17:13 Axel: W3C credentials 10:18:48 harsh: there are many categories in privacy policies (online service, etc.) and there are techniques to extract terms from them (the legal documents). But I'm not familiar with it... does anyone know a bit more? 10:19:04 Axel: I think we need to provide the general framework such that people can plug in 10:19:40 Niklas: People can build upon our initial set of categories 10:21:47 Axel: I would suggest to go to the use cases first and see what is in scope 10:22:40 Rigo: I would suggest to go to the use cases and identify this island. And use vcard first 10:23:06 Axel: But there is an RDF version of vcard and not really new 2011 10:23:18 fajar: There is a version for 2014 10:24:10 ... Interest group note: https://www.w3.org/TR/vcard-rdf/ 10:24:11 vcard ontology: https://www.w3.org/TR/vcard-rdf/ 10:25:00 Axel: Probably it doesn't contain all we might need 10:25:22 Rigo: but that's why we have these island 10:26:00 ... and Linked Data, to find terms that already exists 10:27:13 Javier: We did something similar already with an extension of SPECIAL categories (data, purpose...) for a smart city scenario 10:28:24 (project cityspin: http://cityspin.net/wp-content/uploads/2017/10/D6.1-Privacy-policy-formalization.pdf) 10:28:37 Rigo: And we can use the w3c namespaces for this 10:29:18 ...It is more a social issue rather than a technical issue 10:30:14 Axel: Maybe we should look more at vcard (will add to the wiki) indeed 10:30:42 Harsh: but vcard looks at the data of a user, not the personal data category per se 10:32:08 Rigo: It is easier to look at the real instance data rather than categorize the world. In a second phrase we need to look at the personal data 10:32:54 Axel: We should also consider the time dimension, relation to data and person in the context of the time 10:33:32 Rigo: There is not a standard version to add the context (quads being one of them). There are many, we just need to choose one 10:34:08 Axel: Note that we need to point to the version of the vocabulary being reused (e.g. see at the different foaf versions) 10:35:34 Rigo: our role is to help the "judge" to look at the very fine details. The only think we have to do is to record the version that applies at certain point of time 10:35:41 ... by provenance 10:36:24 Axel: we need to manage the vocabularies in a way that supports versioning (which is fine, we have the knowledge and experience on that) 10:37:08 topic: Check Actions (mostly related to use cases) 10:37:34 rigoo has joined #dpvcg 10:39:12 ... start with Thomson Reuters' use case: https://www.w3.org/community/dpvcg/wiki/SPECIAL/TR_use_case 10:40:09 rrsagent, please set log public 10:40:29 rrsagent, please make minutes v2 10:40:29 I have made the request to generate https://www.w3.org/2018/12/03-dpvcg-minutes.html rigo 10:41:04 Axel: describe the TR use case. Know your Costumer, maximise the info of your costumer in the financial domain (basically for trustworthiness) 10:41:27 ... basically for legal entities, nor physical persons 10:41:59 ... they want to compute the risk level of a legal entity 10:42:23 (going through the types of data in https://www.w3.org/community/dpvcg/wiki/SPECIAL/TR_use_case) 10:44:13 ... same personal document in different regions might have different attributes (e.g., birth document) 10:45:48 rigo has joined #dpvcg 10:47:36 Eva: They are having troubles allocating data to specific sources 10:49:03 Niklas: Note that once you have some basic information, you can derive new data... this is very relevant 10:50:05 Eva: There is article 14th of GDPR that applies when you collect from different sources 10:50:58 Rigo: we are creating an infrastructure for people that want to do the things right, it is a different scenario 10:51:26 Axel: But if collecting public information is subject to GPDR we have to consider it 10:51:27 WEF data categorization only makes sense for blocking tools and other defensive PETs 10:52:39 Eva: sometimes the info of a company is personal information, e.g. if there is a company of only one person 10:53:27 ... also another point to consider in the use case is that, if you have the copy of the passport, you have the full information (not only the ID) 10:53:53 Axel: but if you don't use it, maybe it is not needed to represent 10:54:09 Eva: But data collection matters 10:55:21 Axel: but we can have some "rules", e.g if we have one item of this category, this implies that we have also more information 10:55:52 Rigo: it is really use case dependent 10:57:30 Fajar: what I got from Eva is that it is really different e.g. to fill the ID of a passport in an online form or to have the full picture of the passport as you can have all even if you only need the ID 10:59:05 Rigo: in our use case description (in the website), the description is not complete (with the picture you have more data items potentially) 10:59:46 we are back in categories vs instance data 10:59:50 +q 11:00:49 Axel: what we need then is this kind of implications to know that from an image of a passport you can get X, Y, Z 11:01:00 ack simonstey 11:01:14 -q 11:02:01 Simon: agreed with Axel, if the action is collecting a picture of the passport, then this implies that X, Y, Z information is also collected 11:03:39 q? 11:04:43 Eva: what we need to capture is the kind of information that is relevant when the GDPR applies: processing of personal information 11:05:05 I tried to invite zakim, but he's recalcitrant 11:05:20 ... i.e. operations on personal data (collection,adapting...) 11:05:40 ... collection is also a processing 11:06:40 ... The importan aspect is to understand that GDPR does not only apply when using the data... also collecting 11:06:55 dlewis has joined #dpvcg 11:07:03 s/importan/important 11:07:44 Axel: Agreed, I was suggesting to have this implicit information in the categories, but maybe it is too use case specific (e.g. sometimes the religion is in the passport, sometimes not...) 11:08:41 Bud: For passport, some info is standard (e.g. to read it electronically) 11:09:10 https://en.wikipedia.org/wiki/Identity_document 11:09:40 https://en.wikipedia.org/wiki/National_identity_cards_in_the_European_Economic_Area 11:09:41 Rigo: Bud, it would be good to have these standards 11:09:49 https://en.wikipedia.org/wiki/Public_Register_of_Travel_and_Identity_Documents_Online 11:10:26 Bud: e-ids... it depends on the country 11:10:33 https://www.consilium.europa.eu/prado/en/prado-start-page.html# 11:10:59 Harsh: what about birth certificates? 11:11:46 Axel: not standard, e.g. in Austria you can infer the marital status of the parents 11:12:26 ... maybe the representation is standard, but there are some inferences that are implicit 11:13:25 ... it is a very nice work, to collect the kind of info you can collect from these documents 11:13:30 austrian documents https://www.consilium.europa.eu/prado/en/prado-documents/AUT/index.html 11:14:04 q+ 11:14:06 AxelPolleres has joined #dpvcg 11:14:29 no zakim so far 11:14:44 ... as a CG, it is a good exercise to collect the personal information tht can be inferred from another one 11:15:50 DPVCG-presenter-PC has joined #dpvcg 11:15:53 Simon: I pasted links in IRC (prado coming from EU, an extensive list of possible documents, e.g. weapons cards, driving license...) 11:16:08 we should also take a look at this: https://joinup.ec.europa.eu/sites/default/files/document/2018-11/ISA2%20Study_GDPR%20Data%20Portability%20and%20Core%20Vocabularies_November%202018_1.pdf 11:17:02 Bud: the inference is very difficult, normally is linked to other thinks, and the more I link, the more you can infer 11:18:13 pointer on what might inferred just from people's names from an internationalisation perspective: https://www.w3.org/International/questions/qa-personal-names 11:20:21 stefano has joined #dpvcg 11:20:27 Rigo: From previous work (P3P..) we had three points to standardize: personal data is one of them with sticky policies (link policy with the data), then the object data with processing and what you can do with it, and then the policy and finally the instance data and how to package all together 11:21:36 ... for data we are open, here are some hints (with some core), .... but I understand that we need to define the space of expression that it is given by the GDPR 11:21:50 Niklas: is sensitive data clearly defined in GDPR 11:22:33 Rigo: 2 categories, (i) as defined by article 9, very finite, (ii) and dangerous aspects not mentioned by article 9 such as location data 11:23:36 ... e.g. location data is extremely sensitive (e.g. gender violence) but it is not sensitive within GDPR 11:24:32 Bud: In article 9, there are the conditions when there is a extremely risk.... and I think you can identify location with these conditions 11:24:49 Rigo: Interesting to consider 11:25:08 Bud mentions the Art.29 working Party's Opinion about Data Protection Impact Assessments, they list items and conditions where high risk can be assumed, location being part of it 11:25:50 Axel: So we need the concrete terms in Art 9 + the conditions if we need to categorize sensitive in our taxonomies 11:27:35 Axel: Similarly than the documents, we can collect also some examples of potentially sensitive inferenced data: e.g. if we know that a person visits this place and it is a religious place, then.... 11:28:04 harsh: Do we need to collect some cases of integration, e.g. location+religious=.... 11:28:53 Zakim has joined #dpvcg 11:29:42 Axel: we can maybe first collect the sensitive attributes in our use case. And then see which attributes co-occur and leads to well-known inferred knowledge 11:29:44 present+ 11:29:59 ... to say that there is a potentially risk 11:30:35 s/potentially/potential 11:31:48 present+ Axel_Polleres Eva_Schlehahn Rigo_Wenning Bud_Bruegger, Harshvardhan_Pandit 11:32:01 Bud: It is not only that the attribute is sensitive... In risk assessment in Data protection you have the parties interested, how often you collect, etc. how sensitive something is... is not only about the attribute 11:32:10 meeting: DPVCCG F2F Meeting 11:32:16 chair: Axel Polleres 11:32:34 Fajar: You have to consider the context in order to describe the risk 11:32:57 Axel: Maybe interesting research wise, but for the standard is mostly the attributes themselves 11:33:10 present+ Niklas_Kirchner 11:33:15 present Fajar_Ekaputra Javier_Fernandez Martin_Kurze Amr_Azzam 11:33:23 present+ Fajar_Ekaputra Javier_Fernandez Martin_Kurze Amr_Azzam 11:33:47 AxelPolleres has joined #dpvcg 11:33:50 Zakim, who's here? 11:33:50 Present: simonstey, Axel_Polleres, Eva_Schlehahn, Rigo_Wenning, Bud_Bruegger, Harshvardhan_Pandit, Niklas_Kirchner, Fajar_Ekaputra, Javier_Fernandez, Martin_Kurze, Amr_Azzam 11:33:54 On IRC I see AxelPolleres, Zakim, stefano, DPVCG-presenter-PC, dlewis, rigo, rigoo, Niklas_, Ramisa, Bud, harsh, Amr, Martin_Kurze, Fajar, Eva__ULD_, Javier, RRSAgent, simonstey, 11:33:54 ... Bert, trackbot 11:34:52 AxelPolleres has joined #dpvcg 11:35:03 Dave: There is also a kind of preprocessing, first you contextualise the use case, decide what is personal data and sensitive and then put it in the taxonomy 11:35:36 AxelPolleres has joined #dpvcg 11:37:28 Axel: TR is mostly about the attributes collected in the use case page, let's switch to another use case to have another perspective (deutsche telekom) 11:38:01 harsh has joined #dpvcg 11:38:14 ACTION: Eva to look into requirements of data protection assessment, and whether it would make sense to formalize that in terms of what we standardize 11:38:15 Created ACTION-42 - Look into requirements of data protection assessment, and whether it would make sense to formalize that in terms of what we standardize [on Eva Schlehahn - due 2018-12-10]. 11:40:35 Martin: presenting the use case, slides in the webex channel: https://mit.webex.com/meet/rigo 11:42:28 ... Purpose is (i) to analyse the quality of the network, (ii) to serve a third party, Motion Logic 11:43:45 ... the have the interpolation of the signal, the goal is to come with some traces (e.g. 90% of the people leaving X station go to Y, and 10% to Z) 11:44:22 s/the have/Motion Logic has 11:45:36 ... they can already do it, but they need to prove the quality of their algorithms... and they do it with exact data. That's the main purpose 11:46:19 https://www.motionlogic.de/blog/de/datenschutz/ 11:47:54 ... And that's why derive is so important: you validate your approach with few individuals and you extrapolate to all your customers (millions) 11:48:11 https://www.telekom.com/en/corporate-responsibility/data-protection-data-security/data-protection/your-data-at-dt/details/location-data-492762 11:48:47 Rigo: Besides derivation, it is important to consider that the subject has this control of de-activating the tracking option 11:49:14 AxelPolleres has joined #dpvcg 11:49:53 Martin: collected data is mainly signal strength. But of course with the location data you can infer more things 11:51:33 zakim, who is here? 11:51:35 Present: simonstey, Axel_Polleres, Eva_Schlehahn, Rigo_Wenning, Bud_Bruegger, Harshvardhan_Pandit, Niklas_Kirchner, Fajar_Ekaputra, Javier_Fernandez, Martin_Kurze, Amr_Azzam 11:51:35 On IRC I see AxelPolleres, harsh, Zakim, stefano, DPVCG-presenter-PC, dlewis, rigo, rigoo, Ramisa, Bud, Martin_Kurze, Fajar, Eva__ULD_, Javier, RRSAgent, simonstey, Bert, trackbot 11:51:55 present+ Dave_?? 11:51:56 Martin: the signals stream is linked to other information, such as the device identifier 11:52:22 Dave: Regarding this derivate data... it is a bit probabilistic... is there a threshold to know when derivate data becomes personal? 11:53:03 harsh: that's the filter, you need k-anonymity of 30 people before 11:53:04 Martin: We only use the data if we have at least 30 people in the group, kind of K-anonymity 11:53:21 harsh: still coarse data is personal data? 11:53:34 Martin: the minimum anonymization set is 40 people within a cell tower area, if there are less, the data set is not being used 11:53:37 Axel: location data that is not connected to other data? 11:53:53 Martin_Kurze: can easily be pseudonymized. IMEI to something 11:54:13 AxelPolleres: use case is to derive the model from the sample 11:54:31 Martin: For DT use case, it is more quality analysis and optimization 11:54:49 of the algorithms 11:54:58 harsh: what is the legal basis 11:55:23 Martin_Kurze: it is anonymized and aggregate 11:55:53 Bud: Do you use Android location service? 11:56:00 ... Monetization happens on motion logic side when its not personal anymore 11:56:05 Martin: yes, but it is not implemented now 11:56:15 .. and we only have 30 people in the test, from ML 11:57:55 Harh: Is the Google location service still valid from the consent point of view? 11:58:01 Rigo: not sure.... 12:03:57 we will continue with more use cases after lunch 12:04:06 AxelPolleres has joined #dpvcg 12:04:26 Rigo: please leave the webex now and I will open it again after lunch 12:05:11 will be back at 14h 12:05:16 (1 hour) 13:09:38 AxelPolleres has joined #dpvcg 13:10:42 DPVCG-desktopPC has joined #dpvcg 13:10:52 harsh has joined #dpvcg 13:10:54 we're back from lunch 13:12:29 waiting for a few members to join in the room 13:16:39 Eva has joined #dpvcg 13:18:45 a proposal … as for personal Data categories, we should have a process to propose and accept personal data attributes to the taxonomy: 13:18:46 attribute name: e.g. location 13:18:47 description: physical location of a person 13:18:49 mapping to existing ontologies: 13:18:50 datatype/range: geo coordinates or other geospatial feature identified by a URI (e.g. country) wherein the person is located, that has geo-coordinates or a spatial extent defined by a line or polygon 13:18:52 typical temporal extent: how often is this attribute typically changing for a person 13:18:53 mapping to existing ontologies: e.g. foaf:locatedIn ? 13:18:54 Attribute types: 13:18:55 sensitive according to Art. 9, 13:18:56 requires data protection impact assessment according to Art 29, might depend on an attrribute or combination of attributes at a certain level of (temporal) granularity being potentially sensitive in a certain context 13:18:58 simonstey: can you hear us? 13:19:46 Bud has joined #dpvcg 13:19:50 scribe: harsh 13:20:36 topic: DECODE use case 3 13:20:37 dlewis has joined #dpvcg 13:20:43 Looking at DECODE/DEC03 use-case https://www.w3.org/community/dpvcg/wiki/DECODE/DEC03_use_case 13:22:40 harsh: basically two sets of situations, 13:22:58 ... people living and home and having sensors 13:23:16 https://www.w3.org/community/dpvcg/wiki/DECODE/DEC03_use_case#Alternate_Flows 13:23:27 .... categories of personal data involved, location sensor in the house 13:23:46 .... noise measurement, don't know whether this involves listening devices 13:24:08 Javier has joined #dpvcg 13:24:16 harsh: Do listening devices (from use-case) regard as personal data in the context? 13:24:33 rigo: Listening devices does not come under the GDPR. 13:24:39 (Hasrsh presents the use case https://www.w3.org/community/dpvcg/wiki/DECODE/DEC03_use_case) 13:25:21 rigo: They are considered specifically under telecommunication devices. 13:25:23 q+ to ask rigo whether voice recording, video recording, images are then to be treated differently 13:25:31 ack me 13:25:31 DPVCG-desktopPC, you wanted to ask rigo whether voice recording, video recording, images are then to be treated differently 13:26:09 AxelPolleres: are images of a person personal data? (according to GDPR). If yes, then video would also be. It is weird that audio is not personal data. 13:26:35 rigo: There are rules in addition to the GDPR. The scope of the Telecomm regulations is so narrow that there is nothing of the GDPR left to apply. 13:26:50 rigo: Recordings are very sensitive. 13:27:09 Martin_Kurze: these also apply to devices such as Alexa 13:27:34 rigo: If it indicates noise, then it only indicates presence, and is not sensitive 13:27:52 Eva: it depends if the data is being stored locally or sent anywhere 13:28:02 rigo: This means connected vs disconnected (local) 13:28:52 rigo: If data is being sent elsewhere (remote) then this is of relevance (for purposes of inference) 13:29:31 sensors that transmit their information over electronic communication fall under telecom laws rather than GDPR 13:29:57 AxelPolleres: we never said in the group that we only cover privacy/GDPR 13:30:47 Fajar: Is it different when we share vs when it is always connected 13:31:05 rigo: If the device records data, then it is the same as sharing as it can be shared at a later time 13:31:41 AxelPolleres: is there a difference between audio and video (recording) or still images 13:32:52 rigo: The media type is regulated based on different laws (also case laws) 13:34:01 Eva: What we do for GDPR is also relevant for other contexts (legal), and what is of interest is the processing itself. Whether it is communication. 13:34:17 which is mostly telecommuncation law, penal law, and right on one's own picture 13:34:22 AxelPolleres: We have from the use-case the differentiation between media type (image, video, audio) 13:38:32 what about noise sensors in personal offices... 13:38:37 rigo: If it is a home, and there are multiple people living in the home, then it becomes some form of k-anonymity 13:38:56 Eva: It also depends on whether it is in home or a public place or office. 13:39:08 Rigo: same as k-anonymity, i.e. a moving target 13:40:12 eva, may want to need to define a threashold for those noise recognition devices 13:40:22 and lay down in the ontology 13:44:35 AxelPolleres has joined #dpvcg 13:45:31 we should scope to personal information which does not need deanonymization techniques. 13:46:39 harsh has joined #dpvcg 13:46:49 scribe: harsh 13:46:58 ACTION: eva to have a look a study on AAL that might help us 13:46:58 Created ACTION-43 - Have a look a study on aal that might help us [on Eva Schlehahn - due 2018-12-10]. 13:47:01 the DECODE03 case is a kind of test case for assisted living scenarios 13:48:44 AxelPolleres: I put something in the IRC regarding a proposal for the Data categories - we should start collecting attribute names such as location, dob, a description of what that is (natural language), and whether it already exists in some existing ontology (eg. vcard) and a data type or range (eg. geo-co-ordinates) 13:49:11 AxelPolleres: and the typical temporal extent (eg. how often this changes) 13:49:27 AxelPolleres: this could be the starting point to collect the core attributes 13:50:47 Fajar: Should we only describe the abstract data or also specific data eg. bank account contains other attributes, then what should be included in the data definition? 13:52:50 rigo: this depnds on the use-case. You can have all kinds of data organised in a hierarchy, and if someone links only the top concepts, the structure is lost 13:53:47 dlewis: you start with saying it is a composite object with an anchor (e.g. phone number) and you specify whether it is a strong or a weak anchor 13:54:38 rigo: if we are under GDPR we are obliged to treat the data in a certain way; then it becomes a business decision regarding risk 13:56:23 rigo: it depends on how strong the probability is it that the information can be inferred to be a specific individual 13:56:50 AxelPolleres has joined #dpvcg 13:57:23 AxelPolleres has joined #dpvcg 14:00:57 would it make sense to distinguish directly identifying attributes vs. describing how/under which circumstances this attribute identifies a person? 14:01:49 rigo: its about indentifying identifiers and how they relate via attributes 14:03:13 fajar: should we only add the data/instance e.g. phone number or also model/add information on what can be derived / added to the phone number? 14:04:30 AxelPolleres: we are making this hard for ourselves. In principle, we are looking for data that is one or two hops away from the person (in terms of graph) e.g. bank A/C number 14:04:56 AxelPolleres: attributes also mean chains that can be traced back to the individual, and in the description we say what the chain is 14:05:33 rigo: this is an application of linked data: following the edges 14:05:42 fajar: but then we can run chains of any length 14:06:07 to just one, >1 would be fine? 14:06:38 rigo: it doesn't matter how long the chain is, what matters is that the chain links to the individual. This is the algorithm to use to identify personal attribute for the individual. 14:07:08 fajar: we should not depend on the number of hops / links, it is a technical matter 14:08:04 rigo: to summarise - for personal data, we have a simplified defintion as attributes connected to a person somehow 14:08:23 values of attributes connected (directly or via a known path) to a person. 14:09:04 values of attributes connected (directly or via a known or reconstructable path) to a person. 14:09:26 Eva: regarding pseudo-anonymous data - then someone has the knowledge or path to link it to the individual 14:09:45 rigo: this excludes edge cases, for the moment we only address mainstream use-cases and leave out exotic ones 14:10:21 attributes appearing in a known use case. 14:10:22 rigo: the taxonomy has to be done for personal data first before delving in to more granularity 14:11:35 Eva: (reading definition of persona data from GDPR) 14:12:30 topic: purposes of personal data handling 14:12:38 moving to discussion regarding Purposes 14:13:14 What is a purpose? 14:13:38 rigo: Purpose is what is the policy data - what are you doing with the data 14:14:01 rigo: we should widen the scope to talk about retention time etc. 14:14:12 Javier: definition of personal data by GDPR is very important as it says "any data" related to PII, directly or indirectly (it does not matter if the data can directly identify you) 14:14:37 Eva: purpose is the question "why" so it is the goal for processing - the controller wants to do something 14:14:51 identifiable data is defined in consideration 26 of GDPR: 14:14:53 To determine whether a natural person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person to identify the natural person directly or indirectly. 14:17:16 rigo: in P3P we used "current purpose" and GDPR has reference to this as to fulfil contract or obligation which is an abstract for the current activity. To have a declarative logic there, we have to define all purposes, which are infinite 14:17:44 Eva: we should try to derive classical purposes from the use-cases 14:18:36 AxelPolleres: starting point is purposes from P3P 14:19:10 AxelPolleres: (showing MyData logo) the icons categorise the personal data (maybe in its context) 14:20:08 rigo: these are purposes rather than data categories 14:21:43 harsh: can purposes be structured into a hierarchy? 14:21:47 rigo: most of them yes 14:22:09 (discussion regarding provision of goods as a purpose in an online service, where transaction and delivery are sub-purposes) 14:23:39 AxelPolleres: where would advertisement fit in this (MyData) 14:23:50 rigo: we should add advertising as a purpose 14:24:17 https://github.com/okffi/mydata ... discussion is this a starting point for defining purposes? e.g. where does advertising fall into? 14:25:10 AxelPolleres has joined #dpvcg 14:25:55 rigo: if we have a strict structure or hierarchy then we burden the developers for choosing 14:26:21 eva: we start with a small set of purposes and subsume our use-cases and it is up to others to adopt and apply these approaches for their use-cases 14:26:43 Bud: everyone uses a different vocabulary, and data-subjects want a consistent representation of policy 14:27:19 AxelPolleres has joined #dpvcg 14:28:18 https://www.ownyourdata.eu/en/startseite/ has a simplified classification of the mydata logo 14:28:21 AxelPolleres has joined #dpvcg 14:28:49 AxelPolleres has joined #dpvcg 14:32:06 (discussing SPECIAL categories for purposes in Deliverable 2.1) 14:34:26 Bud: we should express purposes from the perspective of the data subject e.g. storing credit card number for convenience 14:37:50 Eva: shall the purpose reflect the business model / business process where the data is used 14:38:20 we could start with SPECIAL purposes extended by DECODE and smart cities purposes 14:38:44 ACTION: bud to propose/investigate high level purpose classification structuring options 14:38:44 Error finding 'bud'. You can review and register nicknames at . 14:39:35 two goals: Notification and Consent 14:39:39 Eva: purpose should reflect the real part of the business process - what is happening, a connected requirement is that the purpose category should be intelligible to the data subject 14:39:47 purposes could be classified along the beneficiary... e.g. improving of the user experience, vs. optimization of a service in order to make it more cost effective to the provider 14:41:14 rigo: it is a legal matter for being specific vs being abstract; we have to be able to express whatever the outcome i.e. it should be abstract as well as specific in terms of level 14:41:31 fajar: this is a layered taxonomy, and we can provide layers of abstraction 14:41:33 rigo: we wil not be able to solve the granularity problem, that is what courts will have to decide. 14:41:47 fajar: And the organisation can then extend these to be more specific to their services 14:41:58 Eva: we can show this using our use-case (how to extend) 14:42:49 AxelPolleres: whether the purpose is specific or abstract depends on the context e.g. service provision 14:43:17 Eva: We can take one r two of our use cases to show in an exemplary way how the top-level purpose(s) could eventually be sub-categorized to make them more intelligible and precise for data subjects 14:43:24 rigo: service provision already has a legal meaning (from the user's perspective e.g. service of facebook is to communicate socially) 14:43:48 rigo: and then there are 3rd parties that provide advertising which is a completely different purpose 14:46:21 AxelPolleres: so are the purposes sector-based 14:48:25 rigo: we should start with the SPECIAL purposes and use them in the MyData list 14:49:21 education is another purpose 14:53:32 AxelPolleres: how do we structure purpose as definition / taxonomy? It is difficult to express them as a definition using natural language. 14:55:11 rigo: what is the best way to create the taxonomy so that it ensures adoption and reuse? 14:56:04 AxelPolleres has joined #dpvcg 14:56:10 https://www.w3.org/TR/odrl-vocab/#term-purpose 14:59:48 See also some initial purposes in SPECIAL: https://www.specialprivacy.eu/images/documents/SPECIAL_D2.1_M12_V1.0.pdf 15:00:04 I think we had education as well 15:00:14 ok :) 15:00:27 s/ok :)// 15:00:36 svpu:Education 15:03:12 Bud: who benefits from data collection and processing as part of the taxonomy 15:07:59 personal data (collected or inferred) = what? 15:08:11 AxelPolleres: for personal data, processing = how? 15:08:12 purpose = why are these collected? 15:09:07 processing = how are they collected processed to fulfill the purpose 15:20:11 http://bl.ocks.org/susielu/9526340 15:21:08 rigo: how to visualise the purpose hierarchy 15:22:42 or like in the "visualize" tab here https://json-ld.org/playground/ 15:35:28 ACTION: fajar to list purposes from Taxonomy in a table and structure them by source, definition, legal basis (cf. flipchart) 15:35:28 Created ACTION-44 - List purposes from taxonomy in a table and structure them by source, definition, legal basis (cf. flipchart) [on Fajar Ekaputra - due 2018-12-10]. 15:36:39 rrsagent, please draft minutes 15:36:39 I have made the request to generate https://www.w3.org/2018/12/03-dpvcg-minutes.html rigoo 15:36:48 trackbot, close the meeting 15:36:48 Sorry, rigoo, I don't understand 'trackbot, close the meeting'. Please refer to for help. 15:37:08 trackbot, adjourn 15:37:08 Sorry, rigo, I don't understand 'trackbot, adjourn'. Please refer to for help. 15:37:38 trackbot, end meeting 15:37:38 Zakim, list attendees 15:37:38 As of this point the attendees have been simonstey, Axel_Polleres, Eva_Schlehahn, Rigo_Wenning, Bud_Bruegger, Harshvardhan_Pandit, Niklas_Kirchner, Fajar_Ekaputra, 15:37:41 ... Javier_Fernandez, Martin_Kurze, Amr_Azzam, Dave_?? 15:37:46 RRSAgent, please draft minutes 15:37:46 I have made the request to generate https://www.w3.org/2018/12/03-dpvcg-minutes.html trackbot 15:37:47 RRSAgent, bye 15:37:47 I see 4 open action items saved in https://www.w3.org/2018/12/03-dpvcg-actions.rdf : 15:37:47 ACTION: Eva to look into requirements of data protection assessment, and whether it would make sense to formalize that in terms of what we standardize [1] 15:37:47 recorded in https://www.w3.org/2018/12/03-dpvcg-irc#T11-38-14 15:37:47 ACTION: eva to have a look a study on AAL that might help us [2] 15:37:47 recorded in https://www.w3.org/2018/12/03-dpvcg-irc#T13-46-58 15:37:47 ACTION: bud to propose/investigate high level purpose classification structuring options [3] 15:37:47 recorded in https://www.w3.org/2018/12/03-dpvcg-irc#T14-38-44 15:37:47 ACTION: fajar to list purposes from Taxonomy in a table and structure them by source, definition, legal basis (cf. flipchart) [4] 15:37:47 recorded in https://www.w3.org/2018/12/03-dpvcg-irc#T15-35-28