Listen on Spotify
Watch on Youtube
Transcript
Hessie Jones
Welcome everyone to tech and censor. Hello. My name is Hassan Jones. As we see generative AI sore, we saw we see a lot of more. We see a lot more opportunities to advance. Many of our current systems, but we’re also seeing the soaring risks. That are uniquely the result of many of these large language models, so I’m going to throw in some steps when it comes to AI adoption. Mackenzie came out with the their 2024 Global survey on AI. And it indicates that enterprises have adopted generative AI up to 72% this year, and that was up 50% a few years back. Most organizations today that are implementing generative AI will report that it takes. About one to four months to actually put it into production. By 2025. We’re also seeing that 50% of the digital or digital work will be automated and that translates to about 750 million apps that will be using LM’s by that time. And this is massive, but it’s also important to note we’re in the very nascent stage. So that while adoption is growing rapidly, we still have our challenges. So despite the bias and despite the fact that there’s not real knowledge that are happening in these systems, we’re also seeing in accuracy problems, so for example. Insurance companies that use LM’s for their real business data are showing a 22% accuracy when they use them, and this drops down to 0 accuracy for any mid to expert level prompt prompts request. So here’s another this is an interesting. When I saw from Gartner the other day, they predict that 30% of generative AI projects will be abandoned after a proof of concept by, by the end of approval. Concept by 2025. I want to read a quote from Rita Salam who is the analyst at Gartner and she said after last year’s pipe executives are impatient to see returns on generative AI investments. Yet organizations are struggling to prove. And realized value so as the scope of initiatives widen the financial burden of developing and deploying generative AI is increasingly being felt. So what is the cost to actually generate to develop a generative AI? Model to transform your business. It costs between maybe 5 million to $20 million. So many argue that it’s still early day and the effectiveness of these systems will probably happen, but the future viability of of generative AI points to some of the risks. That we’re already seeing and the fact that it’s it’s very different than what we’ve seen in traditional AI machine learning. So when companies like JP Morgan roll out the red carpet and they want to have AI assistance for all their employees. Are they jumping the gun? Is it really the right time to actually go all in in general today? So today, I’m pleased to welcome draw Lee, who is formerly the head of privacy and Data Protection Office at TikTok and is currently the CEO of HYDROX. AI and they offer security and compliance for this new generation of AI. And I’m also pleased to welcome David Danks, who is the professor of data science, philosophy and policy at the University of California, San Diego. He is also a member of the National AI Advisory Committee and also advisor to Hydrops AI. Thank you both. You’re coming.
So today. Thank you, David. Sorry about that. I I talked over you. We’re going to explore this. I guess we called this a paradigm shift because it’s happened almost overnight and anything new that’s come out of generative AI ends up being a unique. Risk. And you know, there’s things that that haven’t even happened yet that that we haven’t evaluated. So those are new things in the horizon that we have to. Consider. So we want to discuss things like that as well as well as the new attack vectors that can be created because of large language models. The demand for data as it increases, how can we make sure that these URL’s become more effective and what are the impacts when it comes to confidential information or personal information? And also the safety of every one of us that that are going. To be using. It so welcome David and and draw and what I’m going to do is I’m going to start off our questions and make. Maybe a little bit more from an educational perspective, I’m going to direct this to David and I’m going to. Ask. When it comes to gender Dai, how does it actually differ from our traditional view of AI machine learning? Can you define it a little bit?
David Danks
It’s always dangerous to ask a professor to define something because who knows where we’ll take it. But I think one of the things that’s really important to understand. About the latest generation of of AI systems, the kinds that really burst onto the public consciousness with ChatGPT, now almost two years ago. Are that they are intended to be much more multi-purpose. I want to use that term instead of general purpose. It isn’t that they can do absolutely everything. I can’t ask check GPT to make me an omelette, but I can’t ask it for a recipe for an omelet. But I can also ask it for tips on how to be a better chess player. Or I can ask it to explain. Complicated biochemical reactions to me, right? So that’s the sense in which these are really multi-purpose. They aren’t constrained to just one task, and that stands in a real contrast to the way that we’ve been doing AI since at least say the early 2000s with the focus very much became. Let’s find one task and build a system that is optimized for that one task and far exceeds human performance at that one task. Whether it’s predicting sepsis in intensive care unit patients or doing image classification as a picture of a dog versus a picture of a cat kinds of things. We’ve really moved away from that, that idea that we’re going to have a bunch. Of little AI. Systems that each do their own specific function and towards these single systems that are capable of many functions in many. Deposits. Now that’s one of the big shifts that. We’ve seen the. Other big shift that we’re seeing with the newest generation of AI systems is in their interface, so it used to be that you would, you know, give a picture and. It would tell. You cat versus dog. Now we’re seeing much more use of what would we would call natural language. Interfaces prompts. So you get to ask a large language model questions and it will respond to you. Or you can ask an image generation system. To give you a picture of an airplane flying through a rainbow filled sky and it will be able. To do that so. That’s another shift in the way that we think about these systems. They’re no longer sitting there with that one task and you put in, you know, you put in your input and outputs and output instead. It’s much more interactive. And much more, driven by the needs and interests of the user, because they can tailor what is coming out of the generative AI system to their needs and interest. Now, having said all of this. I do want to note. That one of the things that’s happened. Over the last two years. Is we’re increasingly seeing people move away from using what? We might think of as sort. Of the the simple or base version of a generative AI system, we’re increasingly seeing people move towards what we often refer to as fine-tuned. AI systems, where they’re specialized to be able to, while still be language centered and multi-purpose. They’re specialized for particular domain or for a particular class. Problems because there does still seem to be a certain amount of generality versus performance tradeoff that if we want to increase performance, we still need to give the system special information. So you know, I don’t want people to think that we are just constantly expanding the scope of what we can do with an AI system. Because in practice, we’re oftentimes narrowing it back down for this particular use cases we have, especially in the business context.
Hessie Jones
OK, so within this new environment, are you seeing or or are you aware of? The risks that continue on from from our traditional AIML perspective and what are the new things that we’re starting to see, especially with the fact that we have these now general models that you could you talked about fine tuning. So in and of that. Perspective. What are the risks when it comes to this?
David Danks
Yeah. So, I mean, we still have many of the old risks that we’ve always had with AI systems where we’ve had for 20 plus years with AI systems, you know, generative AI, large language model, foundation models exhibit really impressive performance, but they still get things wrong. Most notably, what sometimes people refer to as hallucinations, but are really just the system making a mistake. Thinks that a particular book exists, but no book by that name actually exists, or a particular news story exists when it doesn’t actually exist, or, most infamously, court cases. The person who used a a generative AI system to help write their legal brief, and it invented a. Whole bunch of different. Cases and we call these things hallucinations, but they’re really just errors in the way that AI systems have always made errors. And it’s not anything new in that sense. So we have those kinds of old, old, 20 year old problems that we’re used to thinking about. The systems make mistakes, the systems are hard to interpret. We don’t know. We don’t always know why the systems do what they do. But I do think that we’re seeing a new class, a couple of new classes of risks that are showing up as we are seeing the spread of these generative AI, much more multi-purpose Systems 1 is the ways in which the accessibility of these systems has actually increased the likelihood that things are going. To go wrong. Historically, when we would build AI systems, when I think back to some of the systems I was building. A couple of decades ago, we we build systems and then there would be an intensive training process with the user to help them understand. Here’s how you use it. Here’s the kinds of things that are good inputs versus bad inputs, and you could really work with them to make sure that it was used appropriately. Now there’s a prompt based interface. That anyone can access. And that accessibility has led, in some cases, to people doing things maliciously, you know, trying to break the system or use the. System in inappropriate. Ways, But it’s also led to a number. Of just you know. Accidental errors or accidental maliciousness, we might say where somebody because they don’t know how to use this system appropriately. They end up using it in problematic ways and creating new risks. Risks of disinformation, risks of misunderstanding. And so forth. And I think that’s a risk that I, I do believe we’re going to start to overcome back in the early days of Google, Google and AltaVista and search engines. People didn’t know how to use them properly. We had to learn how to do it. And so people were finding the wrong information in those early days and now we’re much better at it. So I do. Think we’re going to overcome that risk. But in the short run, we really have a significant new kind of risk that we haven’t previously had. Now the other class of risks that I wanted to make sure to call out that I think is just inactive here comes along the fact that these models, these generative AI systems are massive. They take an enormous amount of compute power and enormous amount of data, an enormous amount of processing ability, and so there really are only a few companies in the world that are able to create these kinds of systems. And we all know the names, you know, anthropic meta open AI. And and so on Google and.
Really are just a small number of these models. They’re being trained on largely a lot of the same data because they’re all scraping as much data as they can off of the web. And so as a result, we’re seeing the kind of convergence of some of these generative AI systems towards all being kind of similar. And what this raises is a risk of what this was referred to as algorithmic monoculture. The risk is if everybody’s using exactly the same system, then any shortcoming. Or failure in that system is going to have very widespread effects. Now monocultures, this is terminology that comes from agriculture, but monocultures are not intrinsically bad. Most of the you know, servers on the Internet are running all largely the same kinds of code, and we, you know, that’s turning out to be OK. But it’s OK because we put a lot of energy. Effort into safety and security, and we aren’t yet seeing the same level of focus on safety and security, sort of from a monoculture risk perspective. When it comes to generative AI systems. So I think we are seeing some new risks that we’re going to. Have to really think carefully about.
Hessie Jones
OK, I I think at one point when I do speak to draw about this, I I want to talk about like the the algorithm I I think, sorry. Trying to try to protect the algorithms themselves and what does that mean? OK. So so David, when it comes to these new levels of risk, what about the the amount of confidential or private information that’s being scraped and and being input? Into these elops, not to mention. The people that actually within their props may add their own, let’s say, personal information. What what are the implications to that when it comes to the development of these models?
David Danks
You know, one of the big ones, you already called out, which is people, especially when these systems were first being deployed, who would put proprietary company information into a. Prompted to chat, CVT or to Claude and all of a sudden now that’s in the database that open an AI or anthropic, respectively, is training on. So one thing that’s. Already happened is sort of a very rapid I think realization and awareness that we have to be very careful about the information that we. Put into the prompts. Because that’s information that we are now disclosing to some other company. And So what we’re already seeing is a number of organizations that want to take advantage of especially large language models. Who are running their own, who are creating their own in-house large language model so they can control exactly what happens with the data that’s used in prompts or that is used for fine tuning as just one example, we here at UC San Diego have a system that we call Triton, GPT. It’s actually built on the Llama 3 platform from meta. The open source platform, but we’ve deliberately created our own in-house version of this of these technologies, precisely so that we can have. Have confidential information. We aren’t putting student records in there anything. Don’t worry about that. If anyone is listening, but we want to be able to take things. Like you know. The syllabus for my class so that I can have Triton, GPT, help me make better lesson plans or to generate lots of questions for students to test their own knowledge. Before they walk into the room for an exam, we want to be able to have the systems do this, but there’s intellectual property and data ownership questions that come. So what do? We do, we. Build the system here at UC San Diego and we’re lucky enough to have. A supercomputer center here on campus that we can take advantage of to run that. So we’re seeing this sort of shift already happening in under 2 years to organizations realizing that they’re going to have to run their own in order to. Protect their data. I think another. Confidentiality risk that comes up with these systems, though that is perhaps not talked about as. Much is that these systems are incredibly good at synthesizing very large amounts of data, so I can give a large language model 1,000,000 pages of text, and they can very quickly synthesize it and draw patterns and insights from that million pages of text.
A lot of companies and organizations. We’ll have publicly available documents or things that become publicly available in patent filings or other sorts of things where in some. Since some proprietary information is perhaps being leaked through all of these public documents, but in order to do in order to figure out that confidential information, you’d have to read 1,000,000 pages of documentation and who’s going to do that? So historically, we could make certain things publicly available and we didn’t have to worry about. In practice, there being any privacy violations because nobody’s going to go through all of that effort or if they are, they would have tried to get the information through old fashioned espionage. They wouldn’t have bothered with reading everything. Well, that’s all changed with large language models, because now I can go to my competitor and I can scrape everything they’ve ever made publicly available and ask. A large language model to synthesize to extract. What are the what’s the underlying technology that would explain the different patents that they’re doing or where should I look to find out more information about what’s going on at this company that might be proprietary about their processes, practices and technologies. And so I do think that there’s a new kind of threat to privacy and confidentiality. That comes from the capabilities of these systems, not from the data that we are worried about. You know, putting my data in and. Then somebody finds. It out, but whether we have systems. With these incredible. Ability to synthesize information and extract patterns that could absolutely be used to to start finding out information that previously we would have thought would have remained.
Hessie Jones
Can I add? One thing to that and and because Microsoft is, is obviously like the biggest investor in this tech. Knology and, you know, Microsoft Office is used by millions and millions of businesses. Copilot is in there now and there is the chance. And I’m not I’m I’m sure of it that a lot of the chats that are in teams as well. As the the information that you share from person to person, do you think that the the threat of those kinds of information are being are being consumed by Microsoft?
David Danks
So this is this is a very fast moving space. So we are increasingly though seeing organizations that have contracts, whether it’s with open AI or Microsoft or Google, who are making sure that there’s contractual language that says. When those kind, that kind of information can and cannot be used, right? So we saw this happen a few months ago back in the spring zoom, which many of us use basically updated to say well, if you’re using our completely free version, we can, we can take advantage of the content in your zoom call to train our AI. Systems, but basically every corporate client that zoom had had language that said no, they can’t do that right? So we here at UC San Diego used zoom. We didn’t have to worry that our. Conversations were being. Recorded. And so in some sense the. The lawyers of I’m not a lawyer, but the. Lawyers know how to handle this. If they know to ask for it. In the contractual negotiations and that’s where we’re seeing very rapid changes in terms of what language companies are demanding to be put in from. The the user. Side but also from the the technology company. Being able to know what they can and can’t take advantage of as they as you’re saying you know, are trying constantly to find new data. Sources to improve the models.
Hessie Jones
OK, that’s interesting. I could talk about that further and I know LinkedIn actually created a negative option recently on user profiles to say they they basically set it to. No or yes that you would give LinkedIn permission to actually use your information for generative models, but that’s that’s something that we’ll have to. Talk about at. Some point. OK, I want you. I want to talk to you because you understand the value of protecting user information. I mean, you used to work at TikTok. Used to you were you were the head of privacy. So in your new company Hydrox AI, now you’re shifting to maybe not not only that, but really it’s it’s algorithmic attacks on on the large language models, attacks on infrastructure and you so talked a little bit about hydrops. And then explain the areas where data security. Differs today than what we saw before.
Zhuo Li
Right. Of course. First of all, thanks for having me and I’m Joe. I’m the founder and CEO of Hydrax AI. We’re working on a safety and security. So basically we are touching base on a a bunch of things which are quite assist and secure related such as evaluation, right, teaming, basic attack, mitigation, contains monitoring. Model enhancement those. Stuff. So in terms of the question, probably let me quickly talk about how a large language model works and then I can talk about the difference between the the previous or the traditional security versus the the current stage. So you know from very high level, if I just use a text to text generation as an example. So basically there’s a model which which is a huge. Binary that is deployed somewhere where it can handle the traffic, and then there’s a the request comes in, basically a string for example. Can you tell me where I can buy some candies? Right? Something like that. And such a request will be tokenized and mapped to embedded space and from which it will go through a huge amount of the layers where they have a bunch of the nodes and where just trying to activate each of those to generate the next token. And probably if you’re using trade or other models before you will see your responses. Generated one after another. That’s how such a generation works, and those are mostly based on the probabilities. Of course, some different architecture. They may have some different thoughts inside, but mostly that’s how it work, and eventually you get a generated response as a plain text, which answers your question. So normally that’s how it. And probably people can easily tell such a generative AI model operates with a high complex process before producing the results, and this complexity of course, raises questions about how to interpret the output and also how much trust we can play it so which causes or which triggers a probability. Question or the detection question like how much we can understand it, how much we can explain on top of this? So normally, let’s still using a text to text generation example, because image code and modem Daddy is kind of different. So we can make it a little bit simplified. So from there normally when people get a response they barely understand how that generates and from anxiety and security point of view that will cause potential issues for example. Why there’s a his speech, of course. People can mention that there are different algorithms, and those models are trained on different data set and they were they were fine-tuned based on different information and which will cause different results. But of course people can. People probably tried it before that. If you ask the same question multiple times, it’s very possible that you get different answers. Even the meanings are probably the same or similar to each other, so that’s because of the model normally has a balance point between the creativity as well as the how much, how determined the result is. So from that point of view, when we are working on. Air safety and security. The first step we are working on is how to understand this much better before we can run something. So that’s where we have the evaluation at the 1st place. So for example, we are using a model A. Whether this model is safe and secure under which domain under which category and through which attacking method this model is is robust. This model is safe. We need to figure it out before we can move forward. So the first thing is such a transparent understanding is kind of missing. Or there’s no such a benchmark or standardized stuff where everyone can trust or agree upon in place where people can leverage. So that’s the first difference compared with the traditional security space, which are pretty well known and. And that also brings up the second difference, which is the how much we can control it. So from the traditional systems, and we can easily compare like front end, back end database, rational database, Nosql database. But compared with nowadays that’s a giant binary and behind which. It could be like a vector database. It could be red, it could be other things. Those are relatively new stuff versus the previous technology or architecture. So from which? It’s not very guaranteed that we have the full control. So for example, if you’re using vector, database or rag, there’s a X amount of the data X amount of the responses returned, and probably we rank them based on some logic and we picked the top X alongside the prompt that we received and throw that to the model and model will give us some. Sports. Apologies. Yeah. No, no problem. So so from which, uh, we we don’t yet have such a for example, upstream and downstream protection or detection protection at at at this moment and also from the model side, we don’t have enough guard rules based on understanding that we have in place and also for those kind of the models basic data storage for example. Vector database and the rack that I mentioned. We don’t have the full control yet, so that also caused the difference versus the previous version of the traditional security. So basically. Understanding interpretability as well as the control are the things that we don’t have very mature stuff in place which we are actively working on at this moment.
Hessie Jones
It it almost seems like we have this game of whack a mole right now and you’re like the the services that you are putting in place are still learning. To figure out like the the veracity of each of these things and, and I think that that’s where it where I get to. My next question is that the whole thing about interpreting the output, trying to trust you know what that is and the fact that we’re also including. You know scrape data that that as part of these models. And companies are still jumping on the bandwagon like they they want to use it. So from your perspective, because you you deal with these clients all the time, what are their priorities like does? Does it seem like the trustworthiness, the reliability of the models or the data security? Are they really cause for concern at this stage of, let’s say the the? The life cycle. I would say the onboarding life cycle for adoption life life cycle for these companies.
Zhuo Li
Yeah, so from my point of view, and of course, they feel free to try me in that from my point of view and based on the observation and experience in the past one year, each that we are working with the customers as well as we are talking with different. I think for any product the usefulness is the number. One important thing. So if if if there’s a product which is not useful at all and there’s no one is using that, probably people don’t talk much about security and safety, but there are some overlaps. It’s not like you have to be like 100% useful before you start to think. About safety and security. But that has to bring the value on the table that people have people have. Adopted it to some extent and then people start to feel like, OK, I’m going to use this for the long term. And I need. To make sure that safety and security issues are not there. And then I’m going to use it in a more confident way, right? So from that point of view, I think all the customers, they are embracing this, but at the same time, they do have concerns that I’m using something. Like a black box or even it’s like a white box model that I download. I don’t know much about that or I don’t have the full control or understanding of it, so they are trying to find. At this point, like at the same time I want to adopt this more and try to bring value to my business or my team. But at the same time they want to make sure or at least minimize or mitigate those all the known issues at the same time. And there are many companies. I think they are talking about this when they are using the models and of course there are open source models, closed models and also there is a service which is empowered by the models that people can use directly. So from those point of view, I think people are treating those as slightly differently and. There are so many different reasons. For example, if you’re using a service, you just plug and play, and the thing you need to understand is whether or not this service as well as the model behind which empower this service is secure and safe and secure. So that’s where the evaluation is very important. Before people jump on and adopt it. But we’re talking about an A closed model. For example CGT class 3.5. It’s also a service, but it’s empowered by a company and people normally use that. But still it’s like a black box. It’s a little bit hard for people to understand where they can trust this model for their daily business or. In terms of the open source models, it’s a slightly different, but they have different model. I can download it by myself. I will deploy it into my data center. Sounds like I have the full control, but of course there are some other course like I have to understand how to use the model very well. I need to understand how to optimize the influence of the model. Meanwhile, and also whether or not I have the computing power by myself, where I can run this into my business as well so, but of course you have a better control if you if the model is within your data. So basically, those are the trade-offs among different options people are facing right now. And right now there’s no like silver bullet or there’s like a right answer, which is the best choice for the for the users. But I think users they are making trade-offs by themselves. And also at the same time the guidance were a very broadly enforced regulation is not in place yet of course there are so many people working really hard into this topic and trying to make it. Delivered and onboard as early as possible, as in David is working very actively in this space, did a lot of fantastic job, but compared with some like a super mature regulation or guidance or even best practices, it is still new. The whole space is only like 2 years old. Security probably is only like 1 year old and probably the regulation is like a one year old age, so it’s still like very young and there there are there’s a long road ahead of us. But I think until someday that there are the the clear guidance, the best practice where people can confidently leverage and adopt, that’s going to increase people’s confidence where they will use models in different ways, more in their business. And also at the same time that when people are using these models, there are more risks or vulnerabilities coming out every day. Like we are posting a lot of things for sure and also there are other companies who are working on ASP and security. They keep posting about this saying hey I I I identify this and probably they should be fixed by a company or this should be fixed by a service. This should be fixed by a model builder or if you really like this model probably you need to run some fine tuning to make sure that you have the basic. In place before you really integrate this into your product. So those are the things that are happening every day. And I will say that this is just too early stage that a lot of things are mixed together. People are exploring at the same at the same time people are concerning, but I feel like there will be a converge in the very coming future that more things are well established, more things are more transparent and more protection are in place and people will use this more and more.
Hessie Jones
David, did you have anything to add to that?
David Danks
No, I think I think Doug makes. Makes a lot of really good points. I mean, we’re, as you said, we’re still very much in early days and so people are figuring out where they can and cannot trust these systems. I think the vast majority of companies that I know about that are starting to use generative AI systems are basically saying if this is company critical or mission critical, do not assume that it’s right. You know, make sure that a. Human is actually checking it over, which is a way of saying we don’t trust them yet you when you trust things, you don’t have to go back and and check the work after the fact. So I think we’re still figuring out where and when we can trust the systems and it’s going to be a long road. Trust doesn’t happen quickly when it comes to technology.
Hessie Jones
- Thank you. So Zhuo went for your company 1 year old as you said. What? What are the services that your clients are actually using it for?
Zhuo Li
So I can quickly talk about some like very commonly seen and well adopted. Products or modules, so evaluation auditing is definitely one of those, no matter for model builders or model users, they need to have an understanding of what are the issues of their models, no matter they’re building a new version or people are trying to adopt it, they have to have those basic understanding place before they make a decision. That’s number one. #2 is like after you have a model deployed. Of course we need to have the continuous monitoring in place because it’s not like just one off job. Your model is going to serve above of the traffic online on the fly in terms of the influence and a bunch of. Other things so. We cannot always say that, OK, we evaluate this model and. I think this model is 100% secure and safe. We have to have some like online detection and response in place where we can guarantee that something new or something like new attack or some bad traffic comes in. We have the awareness. And the third one I will mention is the right teaming. So right teaming is like not a new term. I mean I think this this was originally from Soviet the right team in translation is called Blue Team or whatever is an absolute way, right team is about attacking. Blue team is about defense or protection. There are two types of the risk. I want to mention here that I will talk about writing. Number one is the embedded or by default issues. For example there’s a model. I don’t do anything wrong, I just normally talk with the model, but the model will give me something bad or something impropriate, so that’s the by default issues, but also probably the model by itself is. Is is OK, but if I do some tricks or if I attack by myself in an aggressive way, the model will start to behave differently. And those will cause some like impropriate image generated or start to give me some wrong information or give me some like sensitive data those are possible. So those are the things we’re talking about as a text and that is categorized under the domain. So for right teaming we are doing two types of the right team at the same time. Number one is for the generic attacking #2 is for the domain specific attacking. Those are relevant but also slightly different. So for example, there’s a very generic model like whatever the model model A and we’re attacking it. So we are trying to bypass the guardrails. So there must be some. Rules in place. We’re trying to bypass those squirrels like jailbreak, prompt injection or prompt. For US type pair, this kind of deal attacking method and for example if we can attack such a model successfully, we are able to gain some benefits or cause some damage to the model. For example, there are some sensitive information of someone’s Social Security number, someone’s physical address, e-mail address. Something if those are part of the training data, there’s a good chance that we can distill those information out that cause the data breach. And also for example, the model doesn’t let me generate some inappropriate pictures, but if I can best bypass the queries they will generate some like definitely inappropriate images or something that will cause misinformation, especially for the election period and for the domestic ones they are. They are quite relevant to the generic. But they are slightly different because they focus on different things, so I can use health care as one example. So health care is one of the domains which has embraced the VDI adoption. Pretty well, I would say. They are using giant UI for notes, taking image recognition for some diseases, automated diagnosis. So for example there are limited amount of the very good doctors but probably they want to scale their capabilities up. So for the very basic diagnosis. If there are enough information, so probably they can let the model to make a decision and tell people why, like why I think in this way and also give them some like medicines. But at the same time, Windows are trained or fine-tuned or integrated with their larger model. There’s a very good chance that this model is being messed up so that the patient or doctors will get some wrong data and that will cause the big issue of mistreated the patient at the same time. Every hospital I think they have some.
Patient information or their secret sources, or there like a past experiences they put those into their data center, or probably originally in the dark format. And of course, when they are using the larger model, they integrate those information into the model and if those data are leaked or if those data are distilled by us. And that will cause the big damage. So at the same time we are doing the right thing. We are trying to focus on not only the generic use cases, but also working with. A couple of the at the very beginning, a couple of the domain special ones trying to focus on the things they care about and they have been using joint VI to help with and we’re trying to attack those to cause issues and identify the vulnerabilities for them to know. And after that we will help them to make mitigate those issues.
Hessie Jones
Yeah, that’s that’s very interesting. And and you mentioned the two sectors, healthcare and finance, highly regulated and you would think that because because of the the sensitivity of the information from those sectors that they actually would be livers. When it comes to adoption and yet that they’re, they’re already, they’re already. Testing it OK.
Hessie Jones
So I have. Two more questions I I want you both to answer the next one because we know a lot of the large language models. That’s great, pretty much all of the Internet and there has been speculation out there that the world is running out of data which is. Fantastically false, I would say at one point we I thought we will continue to make data no matter what, but there is an argument for that and I want to ask your opinion about the use of synthetic data to augment what we have in order to actually create more effective models so. Let’s start with. David. Hi. Sorry. Let’s start with Zhuo and then.
Zhuo Li
So. Either way, but I can speak first. First of all, I think the statement is is partially correct that we are running out of data of the word. I think the accurate way of saying it is we are running out of all the publicly accessible data on the Internet which are easily. Facts by people there are a lot of data which are not publicly accessible. There are a lot of sensitive data which are pretty much private domain and there are a lot of data are generating every day. For example some security attacking methods and relevant data set. We are generating that every day but I don’t think people can get that from anywhere else. So my point of this is. It’s not a bad thing that we’re running out of all the publicly and easily accessible data from the Internet, and that means we hit a milestone where we have those information in place that set up the baseline, that if we have the right amount of the data, if we have the right algorithms, that’s the baseline everyone should hit. Soon or later, but beyond that, there are so many different domains, so many different use cases, landing scenarios, and a bunch of those are probably never. Public, accessible and probably those are forever like privately used information and either people are using those information to train a special model on top of the the the public model, and they’re going to use that in house. For example, the healthcare or the finance we’re talking about. Let’s say that there is an open source models. They are pretty well trained on top of all the public accessible data, but that’s it. That’s probably stage 1 and going to stage 2. There are healthcare and finance and other sectors, education, e-commerce. They must have their own domain knowledge. They must have their own private domain data. So probably there are stage 2. Stage 3, where people are building things on top of the open source model which are well established and on top of which they are making like stronger and more domestic models for their specific use cases. And even going. Further that there are other use cases at the same time, a lot of use cases are probably hardly being satisfied by only the crowd data or probably some data get licensed and that’s where people probably need a synthetic data that to help them with something that they don’t have the. Real world data, but they really need something to. Comprehend or to make their model a little bit stronger to something that they don’t have the data yet, but they really need that. I think that helps with these kind of use cases as a seed. But for the long term, I think it’s going to be the the mixed-use case that synthetic data domain data private data for different use cases. People combine those together. To train the model or fine tune. The model to work. For a a more advanced space in the coming future.
Hessie Jones
Yeah, it’ll be interesting to see whether or not that combination is actually going to lead to efficacy. David, what what are your thoughts on?
David Danks
Yeah, I mean, I I agree that. To the extent that we can talk about the end of data, it’s the end of really easily accessible data that can be done by scraping and you just need a big repository for it. There’s lots of data sources that we that we aren’t really. Using and even if. Even if we just restrict our attention to the easily accessible data, I think there’s a pretty good argument to be made that we are not extracting. The information out of those data, so there’s a lot of work that could be done there, but I wanted to talk specifically about the synthetic data point because I think that, you know this is where a lot of people want to turn. And I think what’s important to realize about synthetic. Data is it’s incredibly useful and powerful when you have pretty good models to generate the synthetic data from right, and you don’t have to know everything that’s going on in your system. The thing that you’re trying to create synthetic data about, you have to know everything about it, but you have to know some things, right? If. One of the places that people have found some interesting successes is in generating synthetic electronic health records. So being able to. Great data that look like a patient, even though they don’t correspond to any actual patients. You don’t have to worry about things like HIPAA here in the United States. The health privacy law. So. But that’s a case where we know a lot, right? We know that nobody is 15 feet tall. We know that nobody IS300. Years old, we. Know that you know you. You typically don’t have. Negative blood pressure, right, we we know a lot of things that can constrain what kind of synthetic data. We generate where I think we run into trouble is when people start talking about synthetic data for things like science or research, where the whole point is we don’t know how the world works. That’s why we’re needing to do the science. Why? We need to. Go and collect the. Data and in those cases I think synthetic data can be very seductive because it looks really easy. It’s a lot easier to generate. Get on my machine here in my office rather than having to go out to the world and collect. It, but I think it’s very likely to lead us down, dead ends or otherwise lead to a kind of negative feedback loop that ultimately results in the model learning to tell whatever the model generated as the data, right. So you get a sort of problematic feedback loop where ultimately the value of the model just collapses because you’re not getting anything new. From it all it’s doing is learning to repeat what we already told it to repeat. So I think it’s there’s a place for synthetic data, but it’s definitely not A1 size fits all solution and there are going to be a lot of problems where it’s just not going to be the appropriate thing for us to.
Hessie Jones
- And I think you made a good point, I research shouldn’t be using synthetic data, especially when you’re developing hypothesis. Maybe we’re talking areas where standards have already been developed and it’s literally just creating a data set out of the things that we already know, right, OK. So one last question, this is a very exciting conversation and I could be here with you guys forever. So David, let’s talk about regulation and we know that right now. Regulation is actually trying to catch up where it was always falling behind when it came to to innovation. We’re seeing things like the UAI Act, which came out in August, the Biden Harris Executive Order, Evening Canada. If you’re aware of the ADA, the artificial Intelligence Data Act. And even as bombs. Are we, like all of this, is starting to come to the forefront and I think government is trying to catch up with understanding. What are the potential harms that LM’s can cause. So from from your perspective, have we come to this inflection point where we we we will now be able to see? A lot more standards and regulations in lockstep with the kinds of outcomes that come from these MLM’s.
David Danks
Yeah. I think as you said, we’re we’re seeing the very rapid emergence of a number of what I would call governance mechanisms. Regulation is one, but there are many others voluntary standards. The kinds of. Safety and security testing that hydroxide does and and other companies are starting to do, can we develop a sort of agreed upon best practices? That’s the sort of thing that we see, for example, in cybersecurity, all the time is a lot of best practices rather than hard regulation in the law because the law changes much more slowly than technology does. So we are seeing the emergence of these governance mechanisms, some of them coming from industry, some coming from government, some coming from academia and civil society. I don’t know if I’d say that an Inflection point yet? Because I think we’re still. Figuring out, as you said, what? All of the risks are, we know, a number of the new risks that arise because of large language models and other. Centered AI systems. But yeah, as we’ve all said a couple of times over the course of this podcast, these systems have really only been in the in the wild, in the public for under two years. And it’s rare that a technology we figure out all the risks and worries in less than two years. I mean, that would be astonishing. If we had. So I think what we’re really moving towards in a lot of places is a kind of iterative loop where we’re trying new governance mechanisms. We’re seeing which ones industry is able to quickly adopt and implement. Then we can look and see. Does this seem to be helping and to the extent that it’s not helping, let’s go back and revisit, let’s update our best practices. Let’s say you know, perhaps decide this is one of these cases where we just are going to need regulation because industry is not going to do it on their own. So the state has to step in and and and. Person. So we’re really we’re entering the the creative innovative iterative phase of governance and regulation right now. And on the one hand that’s very exciting because it means that there’s a lot of possibilities. It means that folks like me at a university can actually influence what’s happening at large companies. Because it isn’t just the province of the lawyers and the the. You know politicians right now at the same time, it’s kind of a terrifying phase to be in because we know we’re going to make some mistakes. We know we’re going to miss some risks and moreover, it’s happening on a global scale. So we’re right now seeing radically different governance ideas. Regulator regulatory ideas show up in the US versus the EU. You mentioned the UAI act. Here in the US, we. Have nothing like an AI act and I’m pessimistic. That we would have one so. We’re doing what we call sector specific regulation. Whereas in Europe they’re doing much more sector general regulation, we have China that is taking a completely different approach from the US or the EU in terms of how they are regulating and governing AI technologies in their borders and companies in a certain sense, don’t respect national borders, right? They’re all of these companies are wanting to operate. Worldwide. And so that’s one of the big challenges I think we have as a governments policy regulatory community right now is how do we start to make the different regulatory governance regimes that are being stood up. How do we start to reconcile them and make them consistent with one another so that the companies, whether hydroxide, I or Google so they know what to plan for, what to design for what the standards are? And so I think in that sense it’s it’s exciting, but very much a roller coaster ride at the moment where it’s not. Quite clear where we’re going. End up.
Hessie Jones
Yeah, I agree. And I I think you’re right about the regulations. It’s going to make it increasingly complicated if we don’t harmonize or standardize some of this stuff because as we know, the minute you develop a technology, it’s not going to be situated in one country, it’s actually going to spread. Of wildfire across different. Jurisdictions and and. Every company will have to play to the rules of that country. Lauren, so thank you so much. I appreciate both of you joining me today. I love the fact that we’re in the midst of change, but like as you said, it’s scary because we don’t know how it’s going to play out, but it’s creating opportunities like hard rocks, AI to be able to figure out how to actually. Create a system that’s going to be here forever, but it has to be effective. So what are the services that it needs in order to ensure that that happens? So thank you. Draw and David for for actually telling us a little bit more about hydroxyl and also what are the implications in the next couple of years as these large language language models actually develop. So for our audience, if you have topics you want us to explore? Please e-mail us at communications at altitudeaccelerator.ca Tech uncensored. It’s powered and produced by altitude accelerator. You can find us on Spotify and wherever you get your podcast. So until next time everyone have fun and stay safe.
Host Information
Hessie Jones is an Author, Strategist, Investor and Data Privacy Practitioner, advocating for human-centred AI, education and the ethical distribution of AI in this era of transformation.
She currently serves as the Innovations Manager at Altitude Accelerator. She provides the necessary support for Altitude Accelerator’s programs including Incubator and Investor Readiness. She will be the liaison among key stakeholders to provide operational support and ultimately drive founder success.
You can also listen to this podcast on Spotify.
Please subscribe to our weekly LinkedIn Live newsletters.