Omar Alonso Interview
A conversation with Omar Alonso, hosted by Alexy Khrabrov. From the FunctionalTV interview archive.
Listen
Download MP3 ↓Follow the words
Read the transcript
[Music] hello everybody I'm Alexey Krylov the hello everybody I'm Alexey Krylov the father an organizer of by Eric this is father an organizer of by Eric this is our first meetup of 2019 we're here on our first meetup of 2019 we're here on location on Domino data labs which we've location on Domino data labs which we've seen grow from a few people at galvanize seen grow from a few people at galvanize the big company now they use Cullen the the big company now they use Cullen the backhand and the serve data scientists backhand and the serve data scientists in the cloud and so we're happy to be in the cloud and so we're happy to be here on location our first speaker is a here on location our first speaker is a moral Lanza whose principal software moral Lanza whose principal software engineer at Microsoft and he is a engineer at Microsoft and he is a veteran by the bands he talked about veteran by the bands he talked about data data by the bay before and so we're data data by the bay before and so we're very happy to have it here thank you very happy to have it here thank you very much romantic saying so can you very much romantic saying so can you tell us a little bit if you know tell us a little bit if you know basically what your focuses and what basically what your focuses and what have you been doing for last couple have you been doing for last couple years years sure so my interest is mostly on the sure so my interest is mostly on the intersection of information retrieval intersection of information retrieval knowledge graphs and human computation knowledge graphs and human computation and we were working obviously with a lot and we were working obviously with a lot of social data in the last two years of social data in the last two years we've been doing a lot of knowledge we've been doing a lot of knowledge graph extraction from different sources graph extraction from different sources not only Twitter what also other not only Twitter what also other Microsoft data sources the talk today is Microsoft data sources the talk today is one of those projects but it's basically one of those projects but it's basically mainly about unsupervised generation of mainly about unsupervised generation of knowledge graphs mmm so are you I think knowledge graphs mmm so are you I think you can't given talks about label you can't given talks about label quality at our previous event and I quality at our previous event and I think that was one of the really think that was one of the really foundational topics so I'm curious how foundational topics so I'm curious how did you come to the area you know like did you come to the area you know like what what is it about data quality we what what is it about data quality we should talk to you one of the reasons is should talk to you one of the reasons is if you want to play with new data sets if you want to play with new data sets like information retrieval and you want like information retrieval and you want to work on search at some point you have to work on search at some point you have to do evaluation and if you don't have to do evaluation and if you don't have ground truth it's very difficult to ground truth it's very difficult to build those data sets and when I was at build those data sets and when I was at a nine I you know learn Mechanical Turk a nine I you know learn Mechanical Turk and I got hooked into the idea of doing and I got hooked into the idea of doing labeling in general that's how I got to labeling in general that's how I got to do labeling that's how by the way I gave do labeling that's how by the way I gave the talk and thanks to you but I'm the talk and thanks to you but I'm trying to finish a book thank selected trying to finish a book thank selected for an invitation because that kind of for an invitation because that kind of we're trying to grow these slides we're trying to grow these slides together in a short monograph which is together in a short monograph which is mostly about the practice of labeling mostly about the practice of labeling mm-hmm there's a lot of publications on mm-hmm there's a lot of publications on how to do labeling but the bulk is how how to do labeling but the bulk is how to do this at scale and involves to do this at scale and involves learning how to design your hits your learning how to design your hits your human intelligent tasks and then based human intelligent tasks and then based on the domain applying different on the domain applying different techniques that's basically the the techniques that's basically the the practice so this is I think you know practice so this is I think you know when I think when you did this talk the when I think when you did this talk the whole like AI wave was not as fine as it whole like AI wave was not as fine as it is to me and now I think what should we is to me and now I think what should we see better relevant examples right see better relevant examples right because the trust in machine learning because the trust in machine learning now I think more people understand and I now I think more people understand and I think there were like mainstream media think there were like mainstream media articles about bias right so if you know articles about bias right so if you know i strangle certain kind of people it i strangle certain kind of people it will not understand all the kind of will not understand all the kind of people their voices right like the people their voices right like the characteristics and so forth so so I'm characteristics and so forth so so I'm wondering kind of do you find that wondering kind of do you find that people understand that the talk comes people understand that the talk comes down to the data originally like did you down to the data originally like did you see more people asking you about this see more people asking you about this like : to amp the data quality oh boy like : to amp the data quality oh boy yes so the answer is yes people are yes so the answer is yes people are getting more interesting and how to how getting more interesting and how to how do I build training sets how do I assess do I build training sets how do I assess the quality of my machine learning the quality of my machine learning models ai models at the same time biases models ai models at the same time biases is and you think that is coming up a lot is and you think that is coming up a lot and last but not least the notion of and last but not least the notion of annotations on labeling datasets because annotations on labeling datasets because sometimes you don't know how you gather sometimes you don't know how you gather this data sets right so I don't even I this data sets right so I don't even I was the sampling methodology which areas was the sampling methodology which areas we touch etcetera etc is usually not we touch etcetera etc is usually not even thought and it's it's been it's even thought and it's it's been it's becoming an crucial because you want becoming an crucial because you want labels you want clean the reset and then labels you want clean the reset and then you not only you want to know how buys you not only you want to know how buys the data studies but also you want to the data studies but also you want to know all the annotations so if you want know all the annotations so if you want to reproduce another experiment so you to reproduce another experiment so you know to refresh a data set you have to know to refresh a data set you have to have all the different knobs that you have all the different knobs that you used to build this data set which in the used to build this data set which in the past was like God gives Alex a data set past was like God gives Alex a data set it was right that's good now it's like it was right that's good now it's like you have to derive these data sets over you have to derive these data sets over and over again for million different and over again for million different properties and domains so you know what properties and domains so you know what comes to mind is rather let's say what's comes to mind is rather let's say what's this called fake news phenomena no right this called fake news phenomena no right obviously news people have different obviously news people have different opinions so if you ask if you want to opinions so if you ask if you want to ask a bunch of people to label political ask a bunch of people to label political tweets Russians will label them very tweets Russians will label them very different than Americans in Russian different than Americans in Russian Americans will label them different from Americans will label them different from either of these groups yes right and so either of these groups yes right and so so I wonder do you find like they do try so I wonder do you find like they do try to compensate for this they want to to compensate for this they want to understand what biases like humans have understand what biases like humans have biases and like do you want to go back biases and like do you want to go back to them and understand what biases they to them and understand what biases they could have had like how do we have to could have had like how do we have to understand biases in humans like do we understand biases in humans like do we have to go back all the way and try to have to go back all the way and try to model I think what you need to you need model I think what you need to you need to do if you're going to build a set of to do if you're going to build a set of sets is to you get an idea if the task sets is to you get an idea if the task has a subjective answer an objective has a subjective answer an objective answer or somewhere in the middle so yes answer or somewhere in the middle so yes if the task is about building objective if the task is about building objective answers then you can get agreements and answers then you can get agreements and compute in traded coefficients etcetera compute in traded coefficients etcetera there are certain things like election there are certain things like election as Democrats versus Republicans you know as Democrats versus Republicans you know fake news versus no fake news those tend fake news versus no fake news those tend to get much harder so it's difficult to to get much harder so it's difficult to get full agreement but what you want to get full agreement but what you want to know is if you sample different times know is if you sample different times the proportions are stable which means the proportions are stable which means like well if 205 thing it's fake always like well if 205 thing it's fake always maybe there's something there buddy one maybe there's something there buddy one sample nobody thinks is fixed on the sample nobody thinks is fixed on the next one everyone thinks it's fake then next one everyone thinks it's fake then you know there's something fishy there you know there's something fishy there right so it's more about trying to find right so it's more about trying to find those tasks were things you're not going those tasks were things you're not going to get agreement on it and then do to get agreement on it and then do something with it versus blindly you something with it versus blindly you know either either blaming the data set know either either blaming the data set or the workers or the interactions which or the workers or the interactions which is usually what is done in practice is usually what is done in practice understanding the nature of the problem understanding the nature of the problem fake news detection is hard mm-hmm fake news detection is hard mm-hmm because like I said if you think in because like I said if you think in terms of US vs. Russian politics is one terms of US vs. Russian politics is one thing but if you say faking news about a thing but if you say faking news about a brexit maybe it's a whole different brexit maybe it's a whole different thing right maybe people make get more thing right maybe people make get more agreement on fake news in Rex's versus agreement on fake news in Rex's versus politics in the u.s. right so politics in the u.s. right so understanding that I think is kind of understanding that I think is kind of the precondition for doing the rest the precondition for doing the rest right and fake maybe I mean right and fake maybe I mean an objective through like something an objective through like something happened and news is wrong then you can happened and news is wrong then you can say it's fake but if something is say it's fake but if something is political opinion right it's political opinion right it's conventionally a question because conventionally a question because somebody speaking is can be blown into somebody speaking is can be blown into some other people I think it's super some other people I think it's super hard crack and also fake for you may hard crack and also fake for you may have a different definition than for me have a different definition than for me maybe you know you find the news that is maybe you know you find the news that is so fake that is hilarious right versus a so fake that is hilarious right versus a lot of persons say in the midwives or in lot of persons say in the midwives or in these girls will get offensive that's these girls will get offensive that's right right so understanding the how humans will so understanding the how humans will answer to the question is this fake or answer to the question is this fake or not it's also very important so you'd not it's also very important so you'd never have to take it for granted and never have to take it for granted and say oh just go and grab some fake news say oh just go and grab some fake news data set because there's always gonna be data set because there's always gonna be very difficult very difficult yeah so you know it's so like that's the yeah so you know it's so like that's the ground true but now you you are into ground true but now you you are into knowledge broths knowledge broths yeah so how do you move from labels to yeah so how do you move from labels to large knowledge drops it's an internal large knowledge drops it's an internal project so I've been working on the project so I've been working on the reason why I got into labels is to reason why I got into labels is to leverage human sensing and with social leverage human sensing and with social data say Twitter or Facebook there's a data say Twitter or Facebook there's a lot of human computation at scale so why lot of human computation at scale so why not just deriving graph from that not just deriving graph from that sensing mm-hmm so if you've already sensing mm-hmm so if you've already stood in today instead of like counting stood in today instead of like counting hashtags and through the topic can you hashtags and through the topic can you derive a graph if you have a lot of been derive a graph if you have a lot of been sharing with veena social network can sharing with veena social network can you derive something else so basically you derive something else so basically aggregations that you can do that are aggregations that you can do that are just beyond counting just beyond counting terms or topics but if you can derive terms or topics but if you can derive these connections has a lot of labeling these connections has a lot of labeling yes Raven is also hard because now you yes Raven is also hard because now you have to label the quality of the graphs have to label the quality of the graphs but it's kind of like the next step on but it's kind of like the next step on on labeling from my perspective in my on labeling from my perspective in my own products I just say so knowledge own products I just say so knowledge graphs is an established area right and graphs is an established area right and so on and if in my mind it was like the so on and if in my mind it was like the what comes to mind the meters are DF and what comes to mind the meters are DF and Semantic Web yes and it's kind of it was Semantic Web yes and it's kind of it was cool for some time yes but not recently cool for some time yes but not recently well I think it was there was a like well I think it was there was a like it's a movie before I die before I kill it's a movie before I die before I kill mother I'm kind of fashionable I write mother I'm kind of fashionable I write there was a notion that like at some there was a notion that like at some point it looked like the bill Semantic point it looked like the bill Semantic Web Web everybody's gonna very like everybody's gonna very like painstakingly underneath their website painstakingly underneath their website protocol and we'll be semantic triples protocol and we'll be semantic triples everywhere and didn't materialize like everywhere and didn't materialize like it's really I think the feeling like it's really I think the feeling like there are some like hardcore veterans there are some like hardcore veterans who still do that right and I feel like who still do that right and I feel like I've seen a lot of work from German I've seen a lot of work from German universities and I think I can see my universities and I think I can see my catalog and everything but but where is catalog and everything but but where is this field I'd like and I met people this field I'd like and I met people from Google from Google I think Eugene Gabriel of which was I think Eugene Gabriel of which was knowledge graphs right so but I think knowledge graphs right so but I think they didn't like it internally to do they didn't like it internally to do some modeling so but I don't see you some modeling so but I don't see you know like mainstream companies talking a know like mainstream companies talking a lot about it like what is it is this lot about it like what is it is this still a very like but I know it's used still a very like but I know it's used internally I just don't know how much internally I just don't know how much with this proprietary how much of it is with this proprietary how much of it is you know open research like what's the you know open research like what's the state of knowledge graphs great question state of knowledge graphs great question so our approach to knowledge graph is so our approach to knowledge graph is not the traditional AI what you just not the traditional AI what you just mentioned it's basically more like mentioned it's basically more like domain-specific graphs which have some domain-specific graphs which have some knowledge on it knowledge on it you mentioned Eugene who is at Google you mentioned Eugene who is at Google Google has has this Google knowledge Google has has this Google knowledge graph graph Microsoft has the Satori knowledge graph Microsoft has the Satori knowledge graph another companies like Amazon they're another companies like Amazon they're willing in all knowledge graph for willing in all knowledge graph for products Pinterest is also building a products Pinterest is also building a knowledge graph for red across the knowledge graph for red across the street so you can think of like little street so you can think of like little little graphs that are specific for the little graphs that are specific for the domains entertainment business etc so domains entertainment business etc so the goal here is not to capture the the goal here is not to capture the world view so that we cannot the psyche world view so that we cannot the psyche and all those projects but it's more and all those projects but it's more like for a specific domain say music you like for a specific domain say music you know who are the black what the bands know who are the black what the bands were the players were the genre etc and were the players were the genre etc and have all these connections if possible have all these connections if possible automatically derive mmm so it's not automatically derive mmm so it's not it's more about a data-driven approach it's more about a data-driven approach to detect entities entity resolution and to detect entities entity resolution and linkage linkage etc then you can expose us a graph with etc then you can expose us a graph with annotation so you can color knowledge annotation so you can color knowledge graph that's how we call it it's not graph that's how we call it it's not like a super you know the world view of like a super you know the world view of although although or the planet but it's a good or the planet but it's a good representation of a specific domain representation of a specific domain interesting so this comes from interesting so this comes from original questionnaires or is it derived original questionnaires or is it derived automatically like what do we need from automatically like what do we need from humans the levels of human rights in the humans the levels of human rights in the case of the teskigi graph that will case of the teskigi graph that will we'll talk today is about detecting we'll talk today is about detecting links people topics detect that and make links people topics detect that and make a few connections you can build some pre a few connections you can build some pre interesting applications you don't have interesting applications you don't have to have you don't have to go full fledge to have you don't have to go full fledge into super crazy techniques which is into super crazy techniques which is understanding people places understanding people places organizations links topics and then make organizations links topics and then make the connections the occurrences or the connections the occurrences or anything else and you can have a pre anything else and you can have a pre decent graph I say graph we could call decent graph I say graph we could call it knowledge graph because have some it knowledge graph because have some knowledge in it but it's not RDF and we knowledge in it but it's not RDF and we don't know we already care on the don't know we already care on the representation of it it's mostly like representation of it it's mostly like which are the different entry points you which are the different entry points you can choreograph and get data around it can choreograph and get data around it that before was difficult in the example that before was difficult in the example of Twitter you can see what people say of Twitter you can see what people say so for the u.s. elections two years ago so for the u.s. elections two years ago before the u.s. elections which is before the u.s. elections which is something is difficult to see today in something is difficult to see today in to her this is just my brother these to her this is just my brother these little little artifacts you see little little artifacts you see interesting so it's kind of interesting so it's kind of domain-specific knowledge drops domain-specific knowledge drops domain-specific knowledge wrap perhaps domain-specific knowledge wrap perhaps is a better it's a better terminology is is a better it's a better terminology is lightweight that let's put it this way lightweight that let's put it this way lightweight cavies which are not super lightweight cavies which are not super comprehensive I mean for the for at an comprehensive I mean for the for at an encyclopedia like Wikipedia but it's encyclopedia like Wikipedia but it's very good at the domain very good at the domain unlike hierarchical or other flat people unlike hierarchical or other flat people on topic so this is one level hierarchy on topic so this is one level hierarchy of topics well which we feel so for the of topics well which we feel so for the one that 1% we don't but if you if you one that 1% we don't but if you if you have the baseline well done then have the baseline well done then aggregations are somewhat easy to aggregations are somewhat easy to develop and we'll show also a couple of develop and we'll show also a couple of examples on those aggregations which is examples on those aggregations which is that the the point of villanies that the the point of villanies domain-specific knowledge graph it's domain-specific knowledge graph it's like the aggregation that you can build like the aggregation that you can build so it's not about the data information so it's not about the data information knowledge so that pyramid is mostly like knowledge so that pyramid is mostly like if you have these little connections if you have these little connections what can you do versus I don't have the what can you do versus I don't have the connections therefore it's difficult to connections therefore it's difficult to do this mmm-hmm interesting so I wasn't do this mmm-hmm interesting so I wasn't gonna shoot gonna shoot a little bit in sense the maybe a little bit in sense the maybe connection maybe not but let's see so so connection maybe not but let's see so so I took to for a socialite who's the I took to for a socialite who's the Kira's creator he's at Google and he's a Kira's creator he's at Google and he's a Google brain right and I talked to him Google brain right and I talked to him about deep learning beginning of this about deep learning beginning of this year last year which was like the peak year last year which was like the peak of kind of deep yearning is gonna solve of kind of deep yearning is gonna solve everything right so basically and he was everything right so basically and he was very interested in the key he kind of very interested in the key he kind of was saying that deploring will never be was saying that deploring will never be general artificial intelligence in general artificial intelligence in Mexico saying it will it's not capable Mexico saying it will it's not capable of generalization because it's very good of generalization because it's very good with pattern detection right but it will with pattern detection right but it will not generalize like if it give you not generalize like if it give you different kind of some different kind of different kind of some different kind of stripes if you give it a completely new stripes if you give it a completely new kind of stripes I shall not you know the kind of stripes I shall not you know the kind of so it will not understand it and kind of so it will not understand it and so he was actually talking about that we so he was actually talking about that we need a fusion of statistical AI and need a fusion of statistical AI and logically I and I think through Russel logically I and I think through Russel Berger talked about that as well so Berger talked about that as well so basically that like the hinting at this basically that like the hinting at this like the limit of statistical AI which like the limit of statistical AI which we are doing right now we are doing right now and they needs to be some combination and they needs to be some combination with what you know cool used to be cool with what you know cool used to be cool traditionally I right so we've got this traditionally I right so we've got this knowledge basis or something else but knowledge basis or something else but basically like I was surprised to find basically like I was surprised to find you know among the kind of cutting-edge you know among the kind of cutting-edge people and deploring the feeling that people and deploring the feeling that you know you need something else to go you know you need something else to go to the next level and and so I wonder to the next level and and so I wonder like the work you do kind of is it kind like the work you do kind of is it kind of the kind of related can it be an of the kind of related can it be an automatic like building knowledge bases automatic like building knowledge bases like basically we need world knowledge like basically we need world knowledge in some way we need something else and in some way we need something else and besides deep learning is this possible besides deep learning is this possible complement to deep learning or it's it's complement to deep learning or it's it's its own kind of specialized area do you its own kind of specialized area do you see any interaction so for the record I see any interaction so for the record I don't know anything about deep learning don't know anything about deep learning so can't comment on that front there one so can't comment on that front there one of the reasons why we've done this of the reasons why we've done this bottom-up and supervised is going back bottom-up and supervised is going back to one of the other points you were to one of the other points you were mentioning is about explained ability so mentioning is about explained ability so if you want to the tag that Alexia it is if you want to the tag that Alexia it is you know related to say dominoes you know related to say dominoes so something how can you explain those so something how can you explain those relationships and sometimes a bottom-up relationships and sometimes a bottom-up approach and I supervise you can explain approach and I supervise you can explain some of these things some of these things mm-hmm that's one of the reasons why the mm-hmm that's one of the reasons why the project that I'll show today or on project that I'll show today or on Twitter is because we want to be able to Twitter is because we want to be able to explain why this thing is related to explain why this thing is related to that and the other thing is provenance that and the other thing is provenance we want to show evidence that these two we want to show evidence that these two things are related and if you can things are related and if you can sprinkle that they are set with sprinkle that they are set with provenance which is a well understood provenance which is a well understood topic and databases then you can make topic and databases then you can make sense of this pursuit of you know this sense of this pursuit of you know this thing is related because it's magic will thing is related because it's magic will be its ability because here's what we be its ability because here's what we believe it's a relationship slightly believe it's a relationship slightly maybe there's a connection I I don't maybe there's a connection I I don't know this one is bottom out that's what know this one is bottom out that's what I'm trying to explain I'm trying to explain I supervise I'm from the ground up I supervise I'm from the ground up grassroots if you want to call it mhm grassroots if you want to call it mhm kind of a lightweight KB's interesting kind of a lightweight KB's interesting yeah I'll let our viewers further yeah I'll let our viewers further comment when they see it cool so so well comment when they see it cool so so well definite looking forward to talk about definite looking forward to talk about like overall where is your research like overall where is your research going and like what are your plans for going and like what are your plans for this year so for this year that was this year so for this year that was gonna be here in San Francisco gonna be here in San Francisco mm-hmm in May I'm running with mm-hmm in May I'm running with co-chairing we love the human co-chairing we love the human computation track which we have some computation track which we have some superb papers the conference is going to superb papers the conference is going to be massive so if you can attend please be massive so if you can attend please attend in October helping out with H attend in October helping out with H comp the human computation covers near comp the human computation covers near the border watching Tuesday the rest we the border watching Tuesday the rest we have a few things or nons graphs that we have a few things or nons graphs that we would like to if we get a forward for would like to if we get a forward for Microsoft to publish that's it we mean Microsoft to publish that's it we mean like under under the covers in the like under under the covers in the trenches right to get you some products trenches right to get you some products that's not a lot of not our research that's not a lot of not our research everything that I've mentioned today has everything that I've mentioned today has been already published and the rest is been already published and the rest is well you know working hard to put it well you know working hard to put it externally at the mall externally at the mall and it's great to hear because you know and it's great to hear because you know we in the community have for instance we in the community have for instance figure-eight for macro flour which do a figure-eight for macro flour which do a lot of human yes can be dangerous I lot of human yes can be dangerous I think this topic is actually kind of think this topic is actually kind of very interested into the motive of very interested into the motive of people in this community so this is people in this community so this is great to have you here looking for their great to have you here looking for their dog thank you for invitations dog thank you for invitations [Music]
Recovered English captions. Automatic transcription may contain errors.
Keep exploring
Follow the guest, their work, and the ideas behind this conversation in the Devreal knowledge graph.
Omar Alonso on Devreal ↗