Saltar al contenido
Let's talk
Back to the blog
Data

Data, Information, Knowledge and Truth

It’s the data, Stupid!

8 min read

It’s the data, Stupid!

About a year ago I designed a talk, or a workshop, called exactly that: "It’s the data, stupid!"

I was using the analogy of Bill Clinton's campaign adviser in the 1992 presidential elections: “it’s the economy, stupid!”. Later on I found out that Bill Gates had used exactly the same creative delirium in another talk. I feel honoured to have unconsciously plagiarised a genius. A clear exercise in cryptomnesia.

The data economy is key today and we clearly live in the era of data. Kai-Fu Lee (VC and author of AI Superpowers, as well as Google's first country Manager in China) referred to this fact by saying: "In the age of AI, where data is the new oil, China is the new Saudi Arabia."

However, data, unlike oil and just like knowledge, is not a scarce resource. Data is unlimited. And the ability to access knowledge is unlimited too. Mahatma Gandhi said "live as if you were to die tomorrow and learn as if you were to live forever". In the era of generative artificial intelligence, data simply feeds models with one basic intention as their objective: to increase the average order value, reduce the cost of maintaining a service, optimise the user experience, allow access to a facility or even design a work of art. Or perhaps multiple objectives depending on the user of the model. Data has become ubiquitous, and so has access to knowledge.

"Live as if you were to die tomorrow and learn as if you were to live forever"–Mahatma Gandhi

Regardless of the value of data as a scarce or unlimited resource. I want to reflect on the idea of data, information, knowledge and truth.

Every April we all consume the content of each new edition of TED. And every year TED opens the doors for us to reflect calmly on the value of trends and to put in sustainable order some ideas worth spreading. This year I have so far been able to watch three valuable talks around the value of generative AI models. Greg Brockman (co-founder of OpenAI) Sal Khan (founder of Khan Academy) and Yejin Choi (professor at the University of Washington)

When we talk about artificial intelligence, we forget a basic concept that underlies any algorithmic model, which is that of augmenting human intelligence. Whether from education, from the perspective of common sense or from the endless source of process automation, this universe of solutions has come to augment our intelligence (and never to replace it). Having an endless source of access to information means increasing knowledge and sometimes not necessarily creating new truths. Perhaps it is an exercise in democratising information and therefore in a global improvement of average knowledge. But does it mean making truth grow?

We already know the answer to this question. Large text models also create confusion and generate hallucinatory delusions. This problem becomes critical when users do not cross-check the information they obtain as a result. Sam Altman, co-founder of Openai.com, has already expressed his fear that these models will be used for disinformation purposes. And it is also true that these models, even at the hands of their creators, deliberately hide information (see the example of ChatGPT hiding that in its GPT3 version more than 60% of its data corpus is over-represented by web pages whose original language is Russian).

We already know that AGI (Artificial Generative Intelligence) is sometimes completely innocent and devoid of common sense (see Yejin Choi's TED talk). And we also know that some models such as ChatGPT generate 7 times more carbon footprint impact than models such as Meta's LLaMA (or at least that is what Meta states in its academic paper). What we do not know is how truth can be improved by this massive access to knowledge.

The first derivative of this is simple enough: if in many models more than 60% of the sources are websites, access to truth seems to be at risk. However, would we say the same when Wikipedia is one of the sources? Do we even know the priority ranking or weighting that has been given to the sources? Can we judge the intellect of the humans who are reinforcing these models?

Too many questions and perhaps too little patience to answer them. Nevertheless, initiatives such as the University of Stanford's Human-Centered AI and the basic principles established by Anthropic (a spin-off of OpenAI and the second most valuable company in the AGI universe) generate incalculable value for making critical and responsible use of this paradigm shift brought about by the massive access to information generated by foundation models.

We are possibly in a moment of delirium about access to information. Perhaps the quantum leap caused by democratising access to knowledge will break the rules of value creation from 2022 onwards. What is also true is that more data is not more information. And it is also true that more knowledge does not necessarily come with greater access to truth. And yes, although there may be as many truths as there are observers, it is also true that some versions of the truth are not free of a clear intention to manipulate our fellow human beings.

Responsible use of this new oracle, in my view, is the only way to move closer to sustainability. We all have people we look up to, human role models to follow in life. We all sense when we are being lied to (even if we find out sooner or later). We all know when the first answer may contain only part of the truth we are looking for. We all know when we are doing something intentionally reprehensible (except for psychopaths and sociopaths, and it seems they are less than 5% of humanity). We can all ask when we sense that we have broken some rule (whether we know it or not). We can all demand that a chatbot reveal to us the information that led it to a particular decision (perhaps with even less embarrassment than asking another human). I think that just as some models already assume that at scale the overall trajectory of AGI use is positive (the biases accumulated by the original data sources and human feedback will be corrected), my long-term view is positive about the impact of humans augmenting their intelligence. In the short term we are going to see cases of ill-intentioned humans making destructive use of such powerful solutions, and also in the short term we will see certain business decisions cutting cognitive capacity in favour of machines (IBM will reduce by 30% its headcount in customer-facing roles by 2025). Nevertheless, in the long term we are facing a quantum leap in the role of humans in the creation of value augmented by ubiquitous access to information. And this transition has to be put in order.

We need to educate ourselves, and fast, and by way of summary we should all do an exercise in reflection and generate a set of basic practices for the responsible use of these models in our day-to-day work. Here is one possible contribution:

  1. Let's learn to push these solutions and teach them to take on the ideal role we would like them to play: [Prompt: “Imagine you are a renowned scientist who is scrupulously careful when citing your sources…]
  2. Let's reinforce their use with unequivocal messages about what is plausible and what is advisable [Prompt: … Avoid using references that do not exist or inventing incorrect sources. Do not lie and do not try to be nice by answering with invented and unverified information]
  3. Let's push these solutions to look after already regulated spaces [Prompt: … In your answer, avoid making use of personal data for which you do not have express and direct consent for the information provided]
  4. Let's build reputation on the responsible use of access to information and on the transparency of the models [Prompt: … Reveal to me the data corpus you are using and give me the latest update level feeding your data sources]
  5. Let's be resourceful and critical in detecting when these models go off on hallucinatory tangents [Prompt: … Don't lie to me, this information I asked you for I obtained with a simple query to this source ]
  6. Let's make the effort to cross-check original sources that we know to be expert in the subject in order to validate the truthfulness of the result offered initially [Prompt: I have checked the original source and I have found the following discrepancies with respect to your previous answer ]
  7. Let's be explicit in our requests and not settle for the first answer [Prompt: … Are you sure this information is correct and does not include some bias or manipulation on the part of its original creator? Can you cross-check alternative sources that hold the opposite view?]

This is a short list of learnings that I promise to keep adding to, and I will be grateful for more contributions from those of you who have tried it. It is not possible to generate a more sustainable world in the era of AI if the intentionality of humans is not sustainable.

“Our intention creates our reality.” – Wayne Dyer.

More Information to keep digging deeper:

The author

Bernardo Crespo

C-suite advisor in AI, data and strategy. CEO of Quantum Markethink and Academic Director at IE. He helps leadership teams make sound decisions in the age of AI.

LinkedIn