SF Text: Jeff Lerman, Q&A with Alexy Khrabrov @Groupon
FunctionalTV interview or Q&A with Jeff Lerman.
Follow the words
Read the transcript
thank you hello everybody I'm Alexi krabrov the hello everybody I'm Alexi krabrov the organizer of SF text and here we are on organizer of SF text and here we are on location on Groupon we have a Meetup location on Groupon we have a Meetup about ontologies and here with us we about ontologies and here with us we have Jeff Lerman staff Anthology have Jeff Lerman staff Anthology engineer at Kaya gym uh who works on uh engineer at Kaya gym uh who works on uh biological uh knowledge organization and biological uh knowledge organization and I'm first curious can you tell us a I'm first curious can you tell us a little bit uh what is the knowledge little bit uh what is the knowledge domain which where you're building this domain which where you're building this ontologies okay so the kyogen ontology ontologies okay so the kyogen ontology which formerly I should say the which formerly I should say the Ingenuity ontology Ingenuity ontology um um covers uh biology broadly is covers uh biology broadly is specifically for use in analyzing specifically for use in analyzing genomic data clinical and research data genomic data clinical and research data as well as more broadly molecular as well as more broadly molecular biology related data so and the original biology related data so and the original impetus for setting something like this impetus for setting something like this up up was probably you know 15 years ago was probably you know 15 years ago as the just the volume of data started as the just the volume of data started to get really large people started to be to get really large people started to be able to do experiments in parallel able to do experiments in parallel where they could acquire a large amount where they could acquire a large amount of data in a short amount of time very of data in a short amount of time very easily easily and the challenge then becomes analysis and the challenge then becomes analysis and so obviously we're still in that and so obviously we're still in that space today where we have a lot of data space today where we have a lot of data and Analysis is is the bottleneck the and Analysis is is the bottleneck the sources of the data are different and sources of the data are different and the volume and the data are different the volume and the data are different but but that's that's why we do what we do so is that's that's why we do what we do so is this anthology of mainly for humans so this anthology of mainly for humans so is it basically organizational this is it basically organizational this knowledge in a human readable form knowledge in a human readable form uh he was supposed to computers right uh he was supposed to computers right but because you can use indulges in but because you can use indulges in different ways sure uh so the ontology different ways sure uh so the ontology is rarely used directly we don't make is rarely used directly we don't make the ontology available we aren't the ontology available we aren't currently currently um uh we don't currently have any um uh we don't currently have any business models we're making the business models we're making the ontology we're directly available to ontology we're directly available to customers uh we do have internally ways customers uh we do have internally ways to visualize what we have and to we have to visualize what we have and to we have our own query language and our our own query language and our manipulation language but the ontology manipulation language but the ontology is really serves as the underpinning uh is really serves as the underpinning uh to other application software that we to other application software that we sell and so uh we need it to be computer sell and so uh we need it to be computer readable certainly computer usable readable certainly computer usable but at the same time it's serving uh but at the same time it's serving uh it's not very far removed from different it's not very far removed from different applications where it has to be applications where it has to be human readable okay okay and so the the human readable okay okay and so the the so the end user is still human it's not so the end user is still human it's not as you know like it's not like you use as you know like it's not like you use this Anthology as a way to automatically this Anthology as a way to automatically cluster things necessarily like you you cluster things necessarily like you you have a human asking queries have a human asking queries right and receiving the data back that's right and receiving the data back that's right sort of yes I mean it's right sort of yes I mean it's interesting the ontology sort of um interesting the ontology sort of um kind of sits at the core of the Hub of kind of sits at the core of the Hub of of what we do because on the one on the of what we do because on the one on the one side we're in the business of one side we're in the business of curating data from uh disparate sources curating data from uh disparate sources from databases they may have a different from databases they may have a different organizational instruction and we've got organizational instruction and we've got um all the which are you know but they um all the which are you know but they still have an organizational structure still have an organizational structure to fully to fully um a sort of open-ended text from from um a sort of open-ended text from from the literature and so we do need the the literature and so we do need the ontology to help the people who are ontology to help the people who are importing that data importing that data um to to assist them in the process of um to to assist them in the process of classification and organization right so classification and organization right so the ontology is both the system by which the ontology is both the system by which we organize data and it informs the we organize data and it informs the tools that we use tools that we use to allow us to have that organized data to allow us to have that organized data and acquire that data on the other side and acquire that data on the other side of course we want to expose that data in of course we want to expose that data in various forms and various forms and various circumstances through our various circumstances through our application software to users and so of application software to users and so of course you know we have human users at course you know we have human users at some point right some point right but the ontology is non-accessed but the ontology is non-accessed directly even by the applications uh we directly even by the applications uh we we take its organizational we take its organizational structure or output we export that in structure or output we export that in some format that is then in turn used by some format that is then in turn used by the application so there's some layers the application so there's some layers in between in between a few layers in between the end users a few layers in between the end users and the ontology itself but it's still and the ontology itself but it's still you know very strongly determined what you know very strongly determined what our applications can do got it so there our applications can do got it so there is a lot of data in in you know the is a lot of data in in you know the people called bioinformatics or genomics people called bioinformatics or genomics and we have actually several startups uh and we have actually several startups uh even as a scholar which I also run and even as a scholar which I also run and uh I was amazed by looking at you know uh I was amazed by looking at you know genomics databases how much text is genomics databases how much text is there how much artifacts so you know there how much artifacts so you know genomic sequences themselves various genomic sequences themselves various papers talking about them various calls papers talking about them various calls right like this paper so and I used to right like this paper so and I used to work at present at the bank a long time work at present at the bank a long time ago and you know they had their own you ago and you know they had their own you know and I looked at Medline and so know and I looked at Medline and so forth and I was really Amazed by the by forth and I was really Amazed by the by the sheer volume of all this data and the sheer volume of all this data and and and so it seems that everybody has and and so it seems that everybody has their own databases their European their own databases their European database that OS databases there are all database that OS databases there are all kinds of uh of things so kinds of uh of things so have this uh that you're talking about have this uh that you're talking about are they unifying all these data sources are they unifying all these data sources for instance you know sequences for instance you know sequences themselves and and papers about themselves and and papers about sequences sequences can he plays both things and you know on can he plays both things and you know on the same Anthology sure so uh you know the same Anthology sure so uh you know we we certainly we we certainly um Define our limits and and so that's um Define our limits and and so that's one thing actually that we don't do in one thing actually that we don't do in the ontology itself is to deal uh with the ontology itself is to deal uh with sequence data sequence data um we do deal with variant data quite a um we do deal with variant data quite a lot and so that's becoming a more and lot and so that's becoming a more and more important as it's become cheaper more important as it's become cheaper and cheaper to sequence uh you know and cheaper to sequence uh you know DNA and and even sequence whole genomes DNA and and even sequence whole genomes um which is sort of just ridiculous that um which is sort of just ridiculous that we're at the point where we can do that we're at the point where we can do that yes uh but uh you know even a single yes uh but uh you know even a single genome results in an enormous amount of genome results in an enormous amount of data even after it's Consolidated and data even after it's Consolidated and the raw data you know much more so uh so the raw data you know much more so uh so uh what's of interest is to be able to uh what's of interest is to be able to look at the differences the the variance look at the differences the the variance uh the genetic variants and to be able uh the genetic variants and to be able to learn something from uh the presence to learn something from uh the presence or absence of different variants so so or absence of different variants so so we don't process whole sequence data we we don't process whole sequence data we don't and we certainly don't store whole don't and we certainly don't store whole papers but we extract papers but we extract um I would say useful information what um I would say useful information what we consider to be useful information we consider to be useful information um statements um statements um more than so qualitative more than um more than so qualitative more than quantitative data statements and quantitative data statements and relationships from relationships from all these sources so from from the all these sources so from from the literature certainly um and there are literature certainly um and there are parts of the literature that we focus on parts of the literature that we focus on very intently very intently um um and then uh there may be there are for and then uh there may be there are for example databases that focus on genetic example databases that focus on genetic variants and we'll import those and what variants and we'll import those and what and so we do have a task of unifying and so we do have a task of unifying um a variety of different data sources so a variety of different data sources so as you mentioned there are American and as you mentioned there are American and European databases for uh for genes and European databases for uh for genes and and other genetic information and other genetic information um I don't know about a competitor to um I don't know about a competitor to the pdb to the protein Data Bank I think the pdb to the protein Data Bank I think that's a fortunately centralized I mean that's a fortunately centralized I mean based on my knowledge from 10 years ago based on my knowledge from 10 years ago right and it's been a while so I also right and it's been a while so I also used to be a protein structural used to be a protein structural biologist and biologist and um but I also haven't thought too much um but I also haven't thought too much about the pdb for a while now but anyway about the pdb for a while now but anyway there are obviously different databases there are obviously different databases covering some of the same information covering some of the same information they complement each other to some they complement each other to some extent and we do our best to unify extent and we do our best to unify information from those databases we want information from those databases we want um you know we can't stand alone uh um you know we can't stand alone uh people coming to us with their data are people coming to us with their data are their data are speaking a certain their data are speaking a certain language right they're using certain language right they're using certain identifiers to refer to Concepts we need identifiers to refer to Concepts we need to be able to map those identifiers to be able to map those identifiers to those same Concepts so we need to be to those same Concepts so we need to be able to take the refseq for example for able to take the refseq for example for one of the uh one of the American one of the uh one of the American database identifiers database identifiers those those IDs and we need to those those IDs and we need to understand those but ditto for the understand those but ditto for the European ones European ones um and that's just for for genes and for um and that's just for for genes and for genetic sequences there's a variety of genetic sequences there's a variety of other things like that other things like that interesting so uh let me give me if uh interesting so uh let me give me if uh throw some idea on the table which I throw some idea on the table which I discussed recently with uh several folks discussed recently with uh several folks in the space so we have you know several in the space so we have you know several startups doing genomics and also amp Lab startups doing genomics and also amp Lab at Berkeley is using spark to to help at Berkeley is using spark to to help basically with sequencing and and basically with sequencing and and essentially there's several there are essentially there's several there are several communities both industrial and several communities both industrial and academic which unite to fight cancer academic which unite to fight cancer through this kind of open genomic through this kind of open genomic research and there is a Global Alliance research and there is a Global Alliance of several forces and so you know being of several forces and so you know being a software engineer myself I'm always a software engineer myself I'm always thinking that you know open source uh thinking that you know open source uh approach is superior because you can you approach is superior because you can you know put something on GitHub and kind of know put something on GitHub and kind of you have a good readme file and good you have a good readme file and good example and test you know uh then then example and test you know uh then then you will have a lot of developers who you will have a lot of developers who will Who will jump on it so so we have will Who will jump on it so so we have these conversations where we can you these conversations where we can you know we're thinking uh how can we know we're thinking uh how can we package this knowledge genomics for package this knowledge genomics for computer scientists and what kind of computer scientists and what kind of resources we need to enable open source resources we need to enable open source collaboration right so we need some kind collaboration right so we need some kind of data which is referenceable and some of data which is referenceable and some common understood format and we also common understood format and we also need to explain the problems uh in need to explain the problems uh in genomics to to developers but because we genomics to to developers but because we have a lot of smart developers in the have a lot of smart developers in the space they prop can probably uh do a lot space they prop can probably uh do a lot of advances so I'm wondering so you're of advances so I'm wondering so you're kind of you're organizing all this kind of you're organizing all this knowledge right and um are there any knowledge right and um are there any insights you can share is the rename insights you can share is the rename and there are any ways for the community and there are any ways for the community to leverage this information and kind of to leverage this information and kind of quickly understand or Point people to quickly understand or Point people to the right pieces right so is this the right pieces right so is this something which which you know software something which which you know software engineer can do as a hobby in their engineer can do as a hobby in their spare time or is this something you know spare time or is this something you know on the professional researcher can do uh on the professional researcher can do uh what is you thinking about this so the what is you thinking about this so the idea here is is you know how much of idea here is is you know how much of what we learn or what we what we what we learn or what we what we organize can be um made available so organize can be um made available so that in a public domain in the public that in a public domain in the public domain or if you know I understand you domain or if you know I understand you guys have a business model and you know guys have a business model and you know we make these products available but how we make these products available but how much of this can be uh can be you know much of this can be uh can be you know you know why by open source Community you know why by open source Community for instance looking at the open for instance looking at the open available date this is something which available date this is something which which Community needs to do uh is this which Community needs to do uh is this you know other ways for Community to you know other ways for Community to organize this knowledge and make it organize this knowledge and make it available to computer scientists kind of available to computer scientists kind of to quickly educate themselves well I to quickly educate themselves well I think yeah I think there's opportunities think yeah I think there's opportunities there there um and frankly I don't think that that's um and frankly I don't think that that's something that we take advantage of or something that we take advantage of or we really promote right now because as we really promote right now because as you pointed out and you know we have you pointed out and you know we have this in common with a lot of people I this in common with a lot of people I think we're a private company and the think we're a private company and the primary uh you know all primary energies primary uh you know all primary energies are devoted towards making things that are devoted towards making things that we can sell unfortunately we can sell unfortunately um on the other hand uh we benefit from um on the other hand uh we benefit from the existence the existence um of public databases of genetic data um of public databases of genetic data and among others I mentioned the and among others I mentioned the genetics that's a big Focus lately but genetics that's a big Focus lately but there are certainly other public there are certainly other public databases databases um that that we use um that that we use um and um and um um you know we do a couple of things so we you know we do a couple of things so we we take advantage of those of those we take advantage of those of those databases but then we also put a lot of databases but then we also put a lot of energy into energy into curating like I said unstructured curating like I said unstructured knowledge and essentially uh knowledge and essentially uh exposing the structure right or exposing exposing the structure right or exposing exposing uh putting putting a structure exposing uh putting putting a structure around the the meanings around the the meanings um um so that ladder stuff is I would say so that ladder stuff is I would say probably going to continue to be pretty probably going to continue to be pretty proprietary because uh that labor proprietary because uh that labor intensive kind of work that just intensive kind of work that just requires a lot of money that's where uh requires a lot of money that's where uh the sort of private domain probably has the sort of private domain probably has an advantage over the public domain an advantage over the public domain um um meaning we know we need a large number meaning we know we need a large number of very highly trained experts doing of very highly trained experts doing this stuff and it's just expensive but this stuff and it's just expensive but uh for the public databases that we use uh for the public databases that we use we um they're more useful to us if we um they're more useful to us if they're higher quality and so we they're higher quality and so we regularly provide feedback to these regularly provide feedback to these databases and they get to know us databases and they get to know us because we because we are so focused on because we because we are so focused on on quality and testing and integration on quality and testing and integration we notice issues that they frequently we notice issues that they frequently don't have the resources to notice and don't have the resources to notice and so we do provide feedback in that way so we do provide feedback in that way and that sort of thing if there were a and that sort of thing if there were a form for it might be even more valuable form for it might be even more valuable if we could say hey we have some if we could say hey we have some suggestions as to how data could be and suggestions as to how data could be and we're happy to share them with the world we're happy to share them with the world because everyone benefits right uh we because everyone benefits right uh we have some suggestions as to have some suggestions as to what people should be paying attention what people should be paying attention to uh when they're developing databases to uh when they're developing databases when they're organizing data when when they're organizing data when they're when they're collating data they're when they're collating data um and uh and here are some techniques um and uh and here are some techniques that people could use uh to that people could use uh to um help that along and to make that um help that along and to make that feasible feasible um so I think that sort of uh um so I think that sort of uh interaction uh I think there's room for interaction uh I think there's room for that sort of thing makes sense so but that sort of thing makes sense so but I'm also curious you know so you said I'm also curious you know so you said you're a biologist right so what brought you're a biologist right so what brought you to from biology to basically you to from biology to basically knowledge organization which is knowledge organization which is essentially you know a computer essentially you know a computer scientist kind of domain it is it is I scientist kind of domain it is it is I mean we're we're an interesting we're an mean we're we're an interesting we're an interesting space uh you know in my interesting space uh you know in my group we we certainly are are doing data group we we certainly are are doing data science as it were and we don't have a science as it were and we don't have a we don't use that name but people people we don't use that name but people people do do um um we are a mix of people with a strong we are a mix of people with a strong computer science background and people computer science background and people with a strong Sciences background I with a strong Sciences background I would say most people in the group are would say most people in the group are PhD level uh scientists either in PhD level uh scientists either in chemistry or biology okay chemistry or biology okay um and with uh varying uh levels of um and with uh varying uh levels of exposure to the clinical side but exposure to the clinical side but probably more on the research side probably more on the research side um and uh it we're basically we're a um and uh it we're basically we're a group of people who are interested in group of people who are interested in um maybe stepping back from the uh what um maybe stepping back from the uh what you might call in the business world the you might call in the business world the individual contributor level when it individual contributor level when it comes to uh scientific knowledge and comes to uh scientific knowledge and more uh sort of aware and interested in more uh sort of aware and interested in the idea that hey there's all this the idea that hey there's all this knowledge out there but it's knowledge out there but it's not that useful if it's not accessible not that useful if it's not accessible and so much of it is inaccessible and so much of it is inaccessible um either because it's not well um either because it's not well um um organized to begin with and that's how a organized to begin with and that's how a structure to begin with OR because it's structure to begin with OR because it's just not just not um integrated into a unified system or um integrated into a unified system or any kind of unified system any kind of unified system and so there's knowledge that can be and so there's knowledge that can be exposed without doing anything in the exposed without doing anything in the lab lab just by sort of putting the pieces just by sort of putting the pieces together that are already out there if together that are already out there if you can reveal those pieces and get them you can reveal those pieces and get them into one system so uh that that's into one system so uh that that's intriguing to me I mean I came from a intriguing to me I mean I came from a place where I was a graduate student and place where I was a graduate student and then a postdoc I guess like we said then a postdoc I guess like we said molecular biology specifically in molecular biology specifically in protein structure and I did sort of protein structure and I did sort of biophysics biochemistry which is a lab I biophysics biochemistry which is a lab I was in the lab I've also but for a long was in the lab I've also but for a long long time I've been interested in long time I've been interested in computers and sort of playing with them computers and sort of playing with them and so that's always been kind of a and so that's always been kind of a of mine and at some point of mine and at some point the idea of using computers as a tool to the idea of using computers as a tool to help organize knowledge and make it more help organize knowledge and make it more accessible and sort of do great things accessible and sort of do great things with it became more appealing maybe than with it became more appealing maybe than the toiling at the lab bench and so I the toiling at the lab bench and so I kind of turned that corner and so I kind of turned that corner and so I haven't it's so important to me to stay haven't it's so important to me to stay connected to the sciences and and to the connected to the sciences and and to the science that I was doing and I'm and I science that I was doing and I'm and I am am um in the you know where we are we're um in the you know where we are we're still thinking very much about the still thinking very much about the science aspect of things science aspect of things um otherwise there's no usefulness to it um otherwise there's no usefulness to it right because it's not abstract right because it's not abstract knowledge like you need to understand knowledge like you need to understand the domain we need to have them really the domain we need to have them really good domain knowledge I mean we need to good domain knowledge I mean we need to understand what we're doing or else and understand what we're doing or else and that's what gives us you know some that's what gives us you know some Advantage Advantage um is that on the one hand we understand um is that on the one hand we understand um how to sort of reduce data and how to um how to sort of reduce data and how to reduce knowledge and model it in a reduce knowledge and model it in a useful way and it'll allow us to do useful way and it'll allow us to do inference and calculation on the other inference and calculation on the other hand we understand the knowledge itself hand we understand the knowledge itself and the bits and pieces and so uh that and the bits and pieces and so uh that hopefully helps us not to make silly hopefully helps us not to make silly mistakes and also to uh sort of mistakes and also to uh sort of recognize the opportunities so I'm going recognize the opportunities so I'm going to kind of come back to this you know to kind of come back to this you know open source uh Community to fight cancer open source uh Community to fight cancer because I was really you know struck by because I was really you know struck by how much uh opportunity is there and how how much uh opportunity is there and how little is uh of this you know little is uh of this you know centralized organization there is a lot centralized organization there is a lot of organizations trying to unify this of organizations trying to unify this data but there is so much like it you data but there is so much like it you know from a computer science scientist know from a computer science scientist point of view is extremely fragmented it point of view is extremely fragmented it also seems to me that you know the also seems to me that you know the result of great Sciences people but result of great Sciences people but they're not necessarily Google level they're not necessarily Google level computer scientists so they because you computer scientists so they because you know they devote their efforts into know they devote their efforts into discovering the primary knowledge and discovering the primary knowledge and there is not enough folks like you who there is not enough folks like you who you know piece it together yet right and you know piece it together yet right and so uh so I'm wondering is it even so uh so I'm wondering is it even feasible is it possible for a computer feasible is it possible for a computer scientists like me who does not have a scientists like me who does not have a formal you know molecular biology formal you know molecular biology training training um but is curious and kind of can um but is curious and kind of can understand almost everything if you read understand almost everything if you read it many times it many times uh presumably uh presumably um Can can we can we present the um Can can we can we present the problems of fighting cancer problems of fighting cancer indigestible pieces can we decompose a indigestible pieces can we decompose a problem and can we also you know take problem and can we also you know take all this knowledge social mapping and all this knowledge social mapping and present a computer scientists and present a computer scientists and digestible pieces and kind of map a digestible pieces and kind of map a little path to them so somebody on over little path to them so somebody on over over a course of time a community can over a course of time a community can kind of make different uh efforts and kind of make different uh efforts and kind of together uh write a lot of good kind of together uh write a lot of good call to eventually you know defeat call to eventually you know defeat cancer through Community you know cancer through Community you know crowdsource programming is it is it a crowdsource programming is it is it a possible possible activity uh right so I possible possible activity uh right so I mean I I think when we when we use terms mean I I think when we when we use terms like defeat cancer we kind of we're at a like defeat cancer we kind of we're at a very high level and and very high level and and um you know sometimes that can um you know sometimes that can you know but but there are lots of uh you know but but there are lots of uh foreign I mean you know so one of the things I mean you know so one of the things people are really interested in doing people are really interested in doing now is uh finding ways to leverage now is uh finding ways to leverage genomic data genetic data uh to genomic data genetic data uh to determine what the best what the most determine what the best what the most likely uh good treatments are for likely uh good treatments are for individual patients yes right so we talk individual patients yes right so we talk about personalized medicine that's a lot about personalized medicine that's a lot of what people are talking about yes of what people are talking about yes um so in order to do that we basically um so in order to do that we basically want to be able to take advantage of want to be able to take advantage of everything that's come before everything that's come before has this variant been seen before if it has this variant been seen before if it hasn't been seen before can we infer hasn't been seen before can we infer things about it based on where it is or things about it based on where it is or or other things that we know so or other things that we know so um I mean there's a few things that um I mean there's a few things that would that would help one um and and would that would help one um and and this is an ongoing Challenge and it's a this is an ongoing Challenge and it's a moving Target moving Target over time we've over time we've we realize and and we when I say we not we realize and and we when I say we not the company but the community realizes the company but the community realizes that uh there might be additional pieces that uh there might be additional pieces of information about a given observation of information about a given observation that are important for us to be able to that are important for us to be able to take full advantage of it but of course take full advantage of it but of course you know if that happens you know if that happens in some year in some year the years you know before that you know the years you know before that you know when people were collecting all that when people were collecting all that data might be missing those pieces data might be missing those pieces um um so there is a certain I think it'll be so there is a certain I think it'll be hard to make a a single unified sort of hard to make a a single unified sort of data format even if you data format even if you which would be great right I mean so one which would be great right I mean so one if we had a public database of all of if we had a public database of all of the relevant observations the relevant observations um that would be a huge step in the um that would be a huge step in the right direction but of course there's a right direction but of course there's a challenge in creating such a thing challenge in creating such a thing because we don't actually know all of because we don't actually know all of the details what are the kinds of detail the details what are the kinds of detail what are the fields I don't know the what are the fields I don't know the kinds of details that we want to we want kinds of details that we want to we want to capture and that's probably going to to capture and that's probably going to change over time so you need a system change over time so you need a system that allows for that flexibility and you that allows for that flexibility and you need approaches that are robust to those need approaches that are robust to those kinds of variations like all right so kinds of variations like all right so some data has some detail and some some data has some detail and some doesn't but we can we still need to be doesn't but we can we still need to be able to do something with it yes uh so able to do something with it yes uh so there's there's those kinds of things I there's there's those kinds of things I think that think that um it's interesting because in a group um it's interesting because in a group like ours uh you know we have this like ours uh you know we have this overlap of expertise between the overlap of expertise between the specific uh genomic and and biological specific uh genomic and and biological domain knowledge and an understanding of domain knowledge and an understanding of data reduction and data organization data reduction and data organization but we're not uh you know world-class but we're not uh you know world-class experts on either side right so experts on either side right so um it would certainly be interesting uh um it would certainly be interesting uh for there to be communication between for there to be communication between groups like ours that are so and there groups like ours that are so and there are other groups that are working are other groups that are working developing biological ontologies or developing biological ontologies or trying to do things like this and people trying to do things like this and people who are who are um um really like 100 data scientists may be really like 100 data scientists may be thinking about some of these problems um thinking about some of these problems um in a more focused way in a more focused way um um and um and sort of have that and um and sort of have that conversation because yes I think that conversation because yes I think that there probably are there probably are um um pieces that we could break off of these pieces that we could break off of these problems problems and have the community sort of turning and have the community sort of turning on them and I think that happens to some on them and I think that happens to some extent but extent but um there could certainly be more yeah so um there could certainly be more yeah so that's I was really motivated by that's I was really motivated by observing all these groups working observing all these groups working towards the same goal but maybe you know towards the same goal but maybe you know as software developers uh we can as software developers uh we can actually facilitate this through open actually facilitate this through open source and it actually strives are source and it actually strives are working on it's probably a very good way working on it's probably a very good way to bring this knowledge together right to bring this knowledge together right so we probably need good public so we probably need good public anthologies as well because if this is anthologies as well because if this is something where you're hanging all these something where you're hanging all these pieces of knowledge onto right so this pieces of knowledge onto right so this is kind of a backbone and you know and is kind of a backbone and you know and so this is something which you know so this is something which you know computer science understand right so you computer science understand right so you know if we have a good public Anthology know if we have a good public Anthology which we all can agree on that probably which we all can agree on that probably will move at least you know a long way will move at least you know a long way so so now I'm really interested in how so so now I'm really interested in how to make that happen yeah so yeah I would to make that happen yeah so yeah I would say that there are um there are attempts say that there are um there are attempts at that out there at that out there um one of the sort of fundamental um one of the sort of fundamental challenges and maybe I'll talk about it challenges and maybe I'll talk about it a little bit later as well of building a a little bit later as well of building a good biology is that it turns out to be good biology is that it turns out to be fairly labor intensive and you can fairly labor intensive and you can certainly with software develop certainly with software develop techniques uh to facilitate maintaining techniques uh to facilitate maintaining uh quality and especially as you you uh quality and especially as you you know increase the domain and increase know increase the domain and increase the coverage the coverage um um but but um the you know the um the you know the what we find are the ontologies that are what we find are the ontologies that are public just don't have that amount of public just don't have that amount of attention being or you know resources attention being or you know resources being devoted to them being devoted to them um either to uh sort of maintain the um either to uh sort of maintain the quality or even to maybe even to test quality or even to maybe even to test the quality right and so at least if we the quality right and so at least if we can get to the testing part we can say can get to the testing part we can say okay here's some good techniques for okay here's some good techniques for testing yes then there could be you know testing yes then there could be you know you know we we would know first of all you know we we would know first of all we'd have a metric of the quality right we'd have a metric of the quality right we would know what the quality was in we would know what the quality was in any given perspective and we would have any given perspective and we would have and people would have things to work on and people would have things to work on they wanted to come and contribute to a they wanted to come and contribute to a project like that it was like well here project like that it was like well here are areas that we know the tests fail in are areas that we know the tests fail in the following ways the following ways so I think that would be a huge step in so I think that would be a huge step in the right direction this I think this the right direction this I think this sounds like something which software sounds like something which software Engineers can really relate to through Engineers can really relate to through test driven development like if first test driven development like if first startable writing tests which is your startable writing tests which is your metric of quality rather than you start metric of quality rather than you start kind of populating and measuring so kind of populating and measuring so that's that strikes me as something that's that strikes me as something issue as a community can probably try issue as a community can probably try but you know thank you very much Jeff but you know thank you very much Jeff it's really interesting and we're it's really interesting and we're looking forward to your talk thank you looking forward to your talk thank you thanks
Recovered English captions. Automatic transcription may contain errors.
Keep exploring
Follow the guest, their work, and the ideas behind this conversation in the Devreal knowledge graph.
Jeff Lerman on Devreal ↗