# Decentralized RAG framework enabling the Verifiable Internet for AI - Branimir Rakic | OriginTrail

- Channel: [ETH Belgrade Community](https://streameth.org/eth-belgrade-community)
- Date: 2024-10-07
- Duration: 27:48
- Topics: People & Blogs
- Watch: https://streameth.org/watch/yt-N0Yicu8ZvNM
- YouTube: https://www.youtube.com/watch?v=N0Yicu8ZvNM

## Transcript

awesome thanks a lot thanks everybody for coming uh I hope um it's been a nice conference for you so far um so um so yeah I'll try to pack in a lot of different topics today um they are a little bit maybe different from your standard um talk at the conference about blockchain so it's very much about Ai and knowledge graphs um and of course blockchain so but yeah to briefly introduce myself I'm one of the founders of origin Trail um and also the CTO um kind of the designer of the decentralized ngraph concept which I'll explain today um originally from Serbia so um yeah we can also later chat in Serbian if you're around and um yeah I'll introduce something called decentralized Rag and it's not rug but rag so um we'll we'll talk about that framework today and what it means for the internet really so I'll just um head over I think I have a bunch of um topics to cover I'll start with what the problem is that we're trying to tackle here so um just as when internet showed up like it was this amazing revolutionary technology very quickly also some anomalies showed up for example one anomaly we all know about is Pam like all of a sudden you had this great connectivity feature everybody could send like an email to anybody and then um it was all nice and great until spam showed up and then everybody there was basically the cost of producing messages went like so low that that an anomaly showed up spam was something nobody wanted or expected uh maybe we could have anticipated it but long story short now many many decades later we have email that have spam filters that work to an extent but actually centralize the architecture if you are an engineer that ever tried to run your own email server server you know what I mean because uh it's a pain in the ass basically you cannot do it like no email provider like Gmail or whatever is going to accept emails from an unknown server even though email was kind of like originally decentralized not in the sense what we see decentralization today um so in way there's kind of a solution but I'm sure nobody in this room is happy with the solution being centralization and like capture by big companies so it's it's an anomaly that like kind of still is not solved and that's spam so what happens with AI so AI brings more more new anomalies that are tricky and they're actually another threat to the internet probably even bigger threat than so far first one you I'm sure you heard of hallucinations I won't explain a lot but basically I'll just explain why because it's relevant for the story later welcome welcome come on in um we're just starting so hallucination that's the thing when you ask like Chad GPT something and it gives you like a response that like is not true it just like burbles something out doesn't have to be Chad GPT it's a feature rather than a bug of llms and um it works be like that because and I'll simplify to the Bone here it's it's not entirely precisely correct but basically llms they are kind of like really good statistical guessers you give it a bunch of words and then it knows what are the best statistically plausible words to come next so they're guessers um and turns out you can do a lot of with a lot of stuff with this guessing but that means also that sometimes they guess wrong very confidently and that's what we call hallucinations so like I said I'm sure you ran into it um You probably since you're here at the conference you have kind of a at least some notion of of uh ownership uh in the digital world uh familiar to you so like in terms of tokens in terms of all kinds of cool things we're seeing in web 3 data ownership is becoming even big even even bigger problem um with AI because um AI just siphons all the all the data everywhere like doesn't care about IP um which was been like a thing that's been happening for for a while in web 2 again centralization today it becomes like even worse um and for example you can see artists in the world they are like basically protesting because of this for a good reason because like you know imagine being an artist who's been like spending your whole life energy into becoming like an actor and then they basically use your your image to create like uh you know new movies with AI without anything that having to do anything with you and I'm sure everybody's like sort of had a little bit of that thought like what happens if AI replaces all of us um so in in a way like we can see the image of that uh actor or actress actually being their data so um that's another problem um big one is centralization like literally I'm sure I shouldn't pitch or talk about this here at some other conferences where which are more AI conferences I spoke speak about this problem a bit more so I'm sure you guys know what this means but basically somebody taking in all the value all the control not a great idea that's not part of our EOS bias is a thing that comes from that to an extent so that means everybody here has has some bias nobody is like perfectly neutral or I have bias um and then Google has bias you can you've seen that with Gemini like creating specifically gender neutral outputs just to try to be more gender neutral when it doesn't even maybe make sense like for example you know Canon Roman Emperor be look like I don't know somebody from India probably not not nothing against people from India or like but it's like you know it's kind of a form of bias rather than hallucination because somebody told the llm do this um and then this is actually probably I'm not sure does anybody know about model collapse here have you guys heard about this problem before okay there's one person you want to tell us in your own words or should I explain it am exactly exactly so really well put basically AI generates data all the time all this new data gets eventually somehow into its input where it learns from that data and then what turns out to happen is that it becomes worse it hallucinates more it like basically and there's kind of mathematical logic to it because what what AI kind of like llms really more like a branch of are doing is they're fitting a curve to some statistical function that represents the world and if you produce a bunch of stuff that sort of messes that curve up because it's a bunch of statistical outputs then of course it's going to try to fit something to something that's no longer probably a sample of the real world so model collapse is already uh been detected with Gro for example and it's actually um considered to be one of the reasons why llms are starting to Plateau which is another opinion it's it's still under research but like basically it's a proven thing that model collapse is is a problem and we expect with all these anomalies we expect that just with as with Spam it was very cheap to create new email it's very cheap to generate new content so pretty much all or like a lot of the content right now is being created with uh with the help of llms um which is not that bad per se uh but it's definitely going to influence this problem the model collapse and uh which in a way in our world of web 3 we can consider as content inflation nobody likes inflation here right so we're all about like 21 million Bitcoins it's never going to change and all these things or like we somehow in governance we control this well content inflation is real happening all the time it's like influencing pretty much everything and causing all of all of these troubles so um what actually I'm trying to going to pitch today is that real content is a new asset class um so but yeah basically if you're really interested into going deep into this we recently released a paper it's a pre-publication super welcome for to to hear your feedback um and that's where we basically talk about this convergence of crypto internet and an AI and we call this idea the verifiable internet for AI you can find it on this link it's about a 10 page document which describes a lot of these Concepts that I'm U going to show today in much more depth um and yeah like I said this is an open idea still very much open for feedback um and contributions so if you find some ideas just let me know uh anywhere Twitter or or whatever Channel telegram um so I'll dive a little deeper into the content of this paper to the presentation just so you can see if it's interesting to go deeper uh but what basically the the the point we're looking at here is to get some sort of emerging collective intelligence in a like perfect world this would be something that that is free of those anomalies that we mentioned at the beginning the problems and we go by um the touring Award winner Yan Leon's uh quote who even though being very opinionated maybe if you follow him you know that he's very very much in a in a fight with Elon Musk right now on Twitter um but um I I think personally and we as a community what what he said here is really true so if we want to build a human collective intelligence system that is um basically um considered to be a a knowledge base of the entire world it has to be crowdsourced so today I'm going to talk about how we're going to crowdsource it um what we're ideally looking for is some of these characteristics so this is the like Dreamland where we want to be we don't want to have hallucinations we want to have uh both data ownership but also information Providence so where does some information come from did I generate it with an LM does it come from like an old school enterprise system is it from your phone that's what we want to know ideally this there's Knowledge crowdsource from multiple places um there's an ability to remain private obviously have some form of data privacy baked in um uh ownership like I mentioned at the beginning and then finally incentives like so that um we break this problem of centralization and and um rather turn the incentives in the system in such a way that it doesn't cause like it did with email as I mentioned at the beginning uh how can we make it so that it um actually promotes decentralization and and more um a more fair value exchange basically so those are some of the characteristics if you look at um kind of a high level diagram of this uh verifiable internet uh for AI idea um as we present a framework for it uh you can um basically Focus on several different layers obviously there's a bunch of data coming in in different models text image video and all kinds of things um and then there's different techniques to use it with AI obviously a lot a lot of these cool things emerging like graph neural networks agents I'm sure you guys heard about that a lot and U but the idea here that we're actually trying to um uh showcase to as many people as possible because it's really powerful is this idea of decentralized Ral augmented generation or de rag like I said at the beginning so um but basically just if you can see in those few bullets what this idea is about is that we have ai operating on verifiable inputs that it's uh basically using this framework the retrial augmented generation I'll explain why that it's uh organized in the decentralized knowledge graph uh so that we can do many different things with with it that it's all based on web infrastructure and the trust route is basically enabled by by blockchains including the incentives um which basically push for positive alignment um so what is rag if you haven't seen rag I'm sure you use chpt pretty much or something of that sort you go you ask CH GPT a question like what is the capital of France and it says Paris and whatever and this is like basically an llm just responding to directly um but like when it starts to hallucinate it's probably because you ask uh ask he asked a little bit more complicated question than that um or it doesn't know about something and usually companies who have some private data for example everybody wants to build their own chat bot on top of their data that's like you know literally everybody is doing it the way they do it is this thing called rag so it's um let's call it conventional even though it's like super new conventional retrieval augmented generation the idea here is you have some UI and you have like an AI model but instead of asking an AI model directly you go and you kind of figure out what data it needs first uh for example you're asking about um I don't know the for example Nvidia earnings report you want to know something from that report it just came out like last week uh obviously Chad GPD doesn't know about it so like it will return some more most likely but if you fetch that data um and you feed it to the model and you say Hey you know find the answer but in that data don't like try to cud hallucinate or figure something else out just like you know read this and chew it up for me and return it that's why it's called retrieval generation because it first does the retrieval in this classic information retrieval sense if you've been like a computer scientist probably you know about this Branch it's called information retrieval it's pretty much all of the search engine work so basically that's the same idea here and then we feed the model and then we tell the model to do generation but doesn't do generation from whatever it knows in its weights we tell it like use that and actually turns out that this minimizes hallucinations quite a bit um you can go into the details very much so we won't we won't have time for that but basically it helps with the hallucination problem yet again it doesn't help with all the problems so like everybody's doing it like I said and everybody's doing it in isolation like each each of the companies they're doing this on their own and we're kind of creating like some form of disconnected uh AI system so um the like I said the big pitch behind the or Trail is that knowledge is a new asset class that this stuff here is actually very valuable and that if we create something called knowledge assets where we package all this knowledge in certain ways um then we're actually able to solve for all of those anomalies um and what these things contain actually a knowledge asset is by the way a live thing you can create a knowledge asset a knowledge in trail right now it contains knowledge in different formats normally usually you would connect some form of symbolic representation which is graph data and neural representation which is Vector data please go ahead go ahead welcome you're coming at the right time you know because the punch line is now yeah take take your spots guys um how much time do we have by the way 10 oh beautiful I have enough time um but I'll with questions I'll speed up then um so basically um you can literally go to the docs and and read much more about this if you're interested but the point I want you to take from here is that the knowledge can sit in any any system uh and long as we wrap it in an nft and provide the according proofs on chain we're able to verify even private knowledge but especially public knowledge and that's why the dkg the decentralized knowledge graph is here um but before I explain the decentralized knowledge graph can anybody tell me but you have to be really quick because we don't have time like a common thread under the hood of all these here do you know like what's their superpower some of you have heard this before the talk because we were having lunch together actually all of them have knowledge graphs under your hood so that means whenever you go to Amazon you want to buy something and Amazon recommends you something else to buy based on your history that comes from Knowledge Graph Netflix recommends you a TV show to watch based on your history comes from an autograph they find similar people who watch things like you and they recommend something Facebook with their posts they basically try to keep you more engaged so you're there on the platform come again Market yeah basically target market point being is you know if you're on Facebook and you're on like if you're not paying Facebook that means you're the product they want to keep you there longer and longer because that creates value for them um and the point here is that they all get a ton of value from putting their knowledge or rather data into knowledge form and there's a distinction between data and knowledge data is very raw knowledge is more like something you would feed to an AI you probably heard of things like annotation there's people who are annotating images like what's on an image so you can train AI uh things like that um well basically knowledge is this higher order than than data it's much more less much more contextualized less raw um and much more connected so they they use it um and they generate a ton of value but they all have closed knowledge graphs you can basically only access kind of something from the Google Knowledge Graph and the reason for that is because their key product is search so and it's basically public so you're able to access so what we're doing at origin Trail is we're creating a decentralized knowledge graph or dkg the architecture is in a most simplified form looks like this there's it's a two- layer system a blockchain layer that supports Parts multiple blockchains and by the way we just we're just also expanding it uh to more more blockchains as we speak Peak basee is the next one that's coming up um it's a separate Data Network that actually hosts a bunch of these knowledge graphs you can run an O on your own um and um it hosts these knowledge assets so the data sits here the vectors the knowledge graphs the nfts the proofs and tokens sit here so we don't overload or blow the state here and and use the chain for what it's not made for it's not made really as a database for data querying rather we use databases that were made for that here and then you can build obviously all kinds of um AI apps on top but the point here being is that the best kind of way to do it is with this decentralized retrieval augmented generation framework But ultimately there's this problem of incentives that I mentioned right so there is a reason why all of these companies built centralized knowledge graphs and the idea of knowledge graphs comes from and some of you maybe heard of it before Tim burners Lee the guy who invented the worldwide web he called it the semantic web web of data was this utopian idea where we connect just like we connect all websites via links we connect all the data on the internet via some form of new links um and it's considered kind of like an idea that never really took off um my opinion is it did but it did in a centralized way so all of these companies they had the resources and knowledge to do such a complex thing um and they got a ton ton of value but the incentive was there they knew how to use it so what if we give incentive in web3 way we have web 3 today it couldn't um have happened before web 3 even existed because even if you look at the architecture that this famous uh inventor worldwide web proposed is that trust is was an element that they kind of added on top at the end um and I would argue it's like it's literally at the bottom it's it's a fundamental element so the idea of this these incentives um um I'll explain in a moment is that we really provide token incent enves to anybody who wants to contribute to such a Global Knowledge base and really crowdsource it so drag is this decentralized retrieval augmented generation what it really does it combines these neural AI elements llms and symbolic AI Elements which are knowledge graphs they are a branch of AI called symbolic AI with decentralized Technologies so essentially llms plus n graphs plus decentralized identity plus incentives tokenization basically on anything you would want to build on a blockchain so you can consider it kind of like this so the dkg is where the knowledge assets live and they live organized in paranet the these paranet are kind of like in in the world of blockchains you can consider them like rollups because they're like different types of things they have different characteristics but in they're not exactly the same they are not like just connected through one single chain together they are more of a kind of like let's say different tables in your database but but it's in the same database so you can query from different tables easily um and an example of a paranet um is um I'm going to show in a moment but basically a paranet is something operated by a paranet operator you can start paranet uh we're literally launching this feature next week on mayet and the idea is you can pick a topic uh let's say ethereum and you can create a knowledge graph a decentralized Knowledge Graph um based on knowledge assets you and your community populate um you can create the AI services on top they can be both onchain or offchain and you can U ask for incentives uh who gives these incentives well the incentives come from a custom chain that the origin Trail Community spawned um called neuroweb and this neuroweb has this neurot token so this process is actually called knowledge mining if you add new knowledge to the network the network rewards you with tokens just like um if you let's say add hash power to the Bitcoin Network Bitcoin um rewards you with Bitcoins here the idea is that you get uh tokens in return for some useful knowledge who determines what's the useful knowledge the paranet operator so this blockchain this neuroweb network basically that's its key feature that it it's it is an evm enabled blockchain it enables all kinds of cool things we're used to but like the core value of it is that it provides neuro incentives for anybody who wants to pitch in knowledge jaes that are valuable who decides what what's valuable well the community um which is by the way 100% uh the the the the the organizer and governor of this chain because it has onchain governance as well an example of an IPO and that is actually happening right now is um by a team called ID Theory maybe you've heard of them they're uh very interested in uh dii so what they want to do is basically they want to collect a bunch of Knowledge from research papers from many in in in history uh also developed in specific ontologies and they want to enable AI agents to produce new um new knowledge or new discoveries where they as a DA maintain ownership through IP on chain uh via these knowledge assets lots of more details to that but I see a nod over there so I'm going to have to finish but I promise I'm at the last slide um so that that was a lot to cover um I'm going to try to suiz how can you join the party so you can basically plug your components if you're a builder you can plug into this idea you can use it on the chain that you're on or like a new chain uh that we can also integrate happy to talk about that um you can build decentralized AI Solutions with drag we have an Inception program where we can also support you through something called uh deliberately chat dkg um so I encourage you going to this website checking it out if you're looking to build something in this area area and see actually how drag works so there's a a live application built on drag um actually running on pulka do and ethereum at the same time uh that you can test um and you can launch your own IPO pick a topic and like basically get get a bunch of incentives from the community finally you can all of this is open permissionless infrastructure so you can run it yourself so you can run the nodes of the DG you can run nodes of neuro web you can stake uh you can contribute to codes so there's a b bunch of ways if this idea seems remotely interesting for you to be um part of it um super happy to talk about that as well later I know that we're out of time I'm just not sure if we're out of time for questions also so I think we can have one quick question one quick question but quick there is one yes here you go yeah thank you for the talk so uh so with the knowledge the thing is for instance if I'm building a model and I need to train it on your data right but I need actually it only once right how would you like I don't know kind of continue supporting this uh uh incentiv incentivization like uh keeping keeping this knowledge actually require constantly required by your users right so actually to keep it like I don't know kind of uh fresh right because for instance if I just like collect data I can like I don't know pay once download everything and actually yeah now I get back to you so yeah it's a bit of a longer question I'll try to rephrase the way I understood it so basically the question is um you let's say as a data buyer somebody who's interested in data you'd say I'm interested in this uh what happens once you download it once and why would somebody keep contributing um so okay it's um there's two scenarios here one scenario is US crowdsourcing knowledge into one shared reusable knowledge base and for that the neurot token is there to incentivize people you can decide what kind of um Quality metrics let's say for knowledge you want to um embed in the system and that's the crowdsourcing part and this basically makes most sense for crowdsourcing public knowledge um what I sense in the question is another case which is knowledge Marketplace features so the ability where you can come and buy some knowledge which is also enabled um and and comes on top so imagine obviously if you're doing some knowledge Marketplace you're doing for some doing that with something which is more like private data it's not going to be public because somebody can just download it and doesn't have to pay for it right so the idea here is that if let's say you find a case where um there's a bunch of private data that people normally won't share I give an example let's say wearables um I don't want everybody to know my like you know bodily functions all the time but maybe you're a buyer for like a crowdsourced amount of like thousands of people giving you this data in an anonymized way um app paret could help people upload their data in an anonymized way uh the public part of the data is actually the part that enables discoverability so when you query the the graph you find aha there's 5,000 entries for you know C rate whatever and then that's what gets uh incentivized with neuro then when you click buy uh when you purchase that data then essentially transaction happens on chain um and that's the second step to it so I hope that gives you a bit of a banner answer and I I apologize that we took so long um it's perfect timing perfect timing thank you so much this than everybody thank you thank you [Applause]
