Introduction to Generative AI and Industry Trends
- Microsoft’s strategic hiring spree highlights the competitive AI landscape.
- AI's rapid evolution is reshaping industries, making AI literacy essential.
- Intellipad offers a beginner-friendly, free comprehensive course covering generative AI essentials.
Two Main AI Learning Paths
- Application path: mastering tools and prompt engineering for practical uses.
- Builder path: deeper focus on machine learning, neural networks, and model training.
- Beginners encouraged to start with applications and gradually explore deeper concepts.
Essential Foundations: Python and Machine Learning
- Python recommended as the primary language for AI development.
- Key libraries: NumPy, pandas for data manipulation; TensorFlow, PyTorch for model training.
- Understanding supervised, unsupervised, and reinforcement learning basics.
Deep Learning and Transformer Models
- Artificial Neural Networks underpin generative AI applications.
- CNNs excel in image tasks; RNNs and advanced versions like LSTM/GRU handle sequential data.
- Transformers, introduced in 2017, revolutionized AI with self-attention mechanisms enabling parallel processing.
- Large Language Models (LLMs) like GPT family leverage transformers for impressive language understanding and generation. For more in-depth information, see the Complete Guide to LangChain Models: Language & Embedding Explained.
Generative Models Beyond Text
- GANs, VAEs, and diffusion models generate images, music, and other creative content.
- Promising tools for creative industries such as digital art and fashion.
Prompt Engineering and API Usage
- Crafting precise instructions (prompts) is crucial for AI effectiveness.
- Mastering context, tone, chaining techniques enhances AI response quality.
- APIs from OpenAI, Google Gemini, and others enable integration of AI into applications.
- To improve skills here, refer to Mastering ChatGPT: From Beginner to Pro in 30 Minutes.
Fine-Tuning and Custom AI Solutions
- Fine-tuning involves training existing models on domain-specific data.
- Tools: Hugging Face Transformers, LoRA for efficient fine-tuning.
- Enables tailored AI applications like legal chatbots, personalized assistants.
Multimodal AI and Advanced Tooling
- AI models that process text, images, audio simultaneously are emerging.
- Platforms like Hugging Face provide pre-trained models and easy deployment.
- LangChain empowers building AI applications with reasoning, tool usage, memory.
- Agentic AI acts autonomously, managing tasks across systems. For clarity on agentic AI distinctions, see Understanding Generative AI, AI Agents, and Agentic AI: Key Differences Explained.
Practical Project Suggestions
- News summarizers, resume writers, image generators using DALLE or Stable Diffusion.
- Multimodal conversational platforms combining speech, text, and images.
- Medical Q&A bots trained on healthcare datasets.
- Deploy projects on GitHub and Hugging Face Spaces for portfolio showcase.
Deep Dive: Understanding Transformers
- Encoder-decoder structure for sequence-to-sequence tasks.
- Attention mechanism computes contextual relevance of each word in a sentence.
- Multi-head attention allows the model to focus on multiple aspects simultaneously.
- Positional encoding adds information about word order.
Open-Source vs. Closed-Source Models and Deployment
- Hugging Face hosts many open-source models enabling research and customization.
- Large models like GPT-4 are typically closed-source and accessed via APIs.
- Enterprise solutions rely on cloud providers (Azure, AWS, GCP) for compliance and data privacy.
- Using API keys securely and managing models within organizational policies is essential.
Retrieval Augmented Generation (RAG) Technique
- RAG combines embeddings-based retrieval from large corpora with generative answering.
- Process:
- Embed user query.
- Compute similarity with document embeddings.
- Retrieve top relevant chunks.
- Pass retrieved context plus question to LLM to generate accurate answers.
- Enhances response accuracy and handles large knowledge bases.
LangChain: Simplifying AI Application Development
- LangChain provides abstractions for document loading, indexing, retrieval, and prompt management.
- Supports integration with multiple data sources, vector stores, and LLMs.
- Enables constructing complex workflows with chaining and agentic capabilities.
- Example usage includes web scraping, document chunking, vector indexing, similarity search, and answer generation.
- For foundational concepts and alternatives, see Understanding LangChain: Importance, Applications, and Alternatives.
Advanced Prompting Techniques
- Few-shot learning: providing examples within prompts for improved model responses.
- Chain-of-thought prompting: encouraging step-by-step reasoning for complex problem-solving, especially math.
- Importance of crafting prompts to control output format, tone, and factual accuracy.
Summary
- Generative AI today combines foundational neural architectures with vast datasets and advanced training techniques.
- Practical AI development involves mastering prompt engineering, APIs, fine-tuning, and retrieval systems.
- Tools like Hugging Face and LangChain make building AI applications accessible and scalable.
- Staying updated and skilled in these areas unlocks career opportunities in the fast-growing AI industry.
For a full course on generative AI and certification, visit the Intellipad program powered by iHub IIT Roorkee described in the video.
Just when we thought the AI race couldn't get any crazier, Microsoft made a silent yet powerful move. Last week,
they hired over 20 top AI engineers from Google Deep Mind without much noise, but with huge impact. One of the most talked
about highest, Wun Moan, the brain behind the startup Windsor, which was recently acquired by Google in a $2.4
billion deal. We are not just watching a trend but we are witnessing an AI battlefield where Microsoft, Google and
startups are fighting for the minds that will shape the next era of intelligence. Why is all this happening? Because
everyone wants a piece of AI future from chatbot to enterprise AI tools. Every company's racing to build smarter,
faster, more humanlike technology. And here's the thing, you don't have to be a tech giant to be a part of it. But if
you're sitting on sidelines, you're already a step behind. That's exactly why we at Intellipath have created the
most practical and beginner friendly genai full course absolutely free. We have broken down everything you need to
know from deep learning algorithm, genai models, transformers, autoenccoders to hands-on tool like lang chain, hugging
face, MCP servers and even building your own AI agent. This video is your one-stop destination to confidently
start your journey into generative AI in 2025. So take your laptop, tune into Google Collab, and let's dive deep into
an immersive Gen AI learning experience right here on Intellipad's YouTube channel. Our tech revolution has already
begun. Just look around. Genai hiring, Gen AI is in picture. The AI age is here. We have reached a point in history
where we can build an app without writing a single line of code, create art without picking up a brush and write
a script, design a product, launch a business just by giving instructions to AI. Generative AI is growing fast. The
industry is worth over $ 38 billion in 2025 and is expected to cross $1 trillion in less than a decade and
companies are already hiring for roles like generative AI engineer and prompt engineer. But while these roles are
emerging, thousands of jobs are also disappearing. The layoffs are real and they are hitting hard to lay off over
12,000 job roles altogether. This time there is a clear culprit. It's artificial intelligence. If you're
wondering whether AI is coming for your job, well, spoiler alert, it may already have.
>> People in tech, marketing, design, and customer services are losing job. Not because they are not talented, but
because the tools and industry have evolved. The hard truth, skills that were valuable 5 years ago aren't enough
anymore. If you're not adapting, you are at risk of becoming replaceable, not by a person, but by a tool. And that's
exactly why this video matters. There's a small window of opportunity right now where anyone who decides to learn and
adapt can actually lead this stage. You don't need to be a coding expert or graduate from a top college. You just
need the right direction. And that's what I'm here to give you. Presenting the complete generative AI road map. A
simple 10-step guide for absolute beginners. Whether you are student figuring out your path, a working
professional trying to stay relevant, or someone genuinely excited about AI, this road map is your starting point. This
road map if you follow and study as discussed in the video you will be able to crack generative AI roles or rather
build your own genai product down the line. You can find the road map in the description below for absolutely free.
So let me clear the air by explaining two different path you can take to be a geni pro. Let's look at the very first
step. Understand the two genai path. Before we jump into coding or training model, it's important to understand
where you are headed in generative AI. There are two main routes. The first is the application path. This means using
genai tools smartly. You'll learn how to write effective prompt, use tools like chart GPT or DL and integrate AI into
real world app using APIs. The second path is for builders. Here you go deeper into how AI works behind the scene. You
learn machine learning, neural network and transformers. Basically, how these models are created and train. Most
people start on application path and slowly build the confidence to go deeper. So don't worry if you are a
beginner. The key is to just begin. Step two, learn a programming language. See, you can either go for JavaScript or
Python. But I would recommend you to learn Python. Python is the language that powers almost all AI development
today. If you have never coded before, don't worry. Python is bigger friendly. You can learn the basics like loop,
function, and condition within a few weeks. Once you have got the basics, move on to two essential libraries which
is numpy and panda. These are essential for working with data and nai because data is everything. Numpy helps you with
numbers and array while panda help you load and clean data from files like CSVs. See the coding you need to work
around AI is not just build app rather it's more about training models using available frameworks and libraries like
TensorFlow, PyTorch, third party APIs and more. You can learn Python from Google Python class, Python's official
documentation. We ourselves have recently rolled out machine learning course. You can check it via the link in
the description. Step three, learn machine learning. Now that you're comfortable with Python, let's answer
the big question. How does a machine actually learn and generate? This is where machine learning or ML comes in.
Imagine you give a machine thousands of example like houses with their size, location, and prices. Over time, the
machine starts to recognize patterns such as houses in location X with 2,000 square ft usually cost around this much.
It doesn't memorize, it learns from pattern in the data. That's the magic of ML. Start by learning main types of
machine learning such as supervised learning. You give the machine board the input like email and target variable.
Meaning let's say you want machine to predict price of house on basis of location, carpet area, number of rooms
etc. Then initially you also give the price as well for training. After training machine would be able to
understand the pattern and predict the house prices for new entries you make. This is basically a simple explanation
of how supervised learning works. Common model includes linear regression, logistic regression, decision tree,
random forest, SVM, KN&N. Moving on to unsupervised learning. In simple word, it's when you give your model a bunch of
data, but you don't label to figure out the pattern. The model has to figure things out on its own and come up with
pattern detection, grouping, etc. Now, let me give you a real world example to make it more clearer. Imagine you run an
online store. You have tons of customer data. how much they spend, how often they visit, what type of product they
buy, their age, where they live, and so on. But here's the thing, you don't know whether you should retarget them by
advertisement or if they're already a loyalist. This is where unsupervised learning comes in. You use a technique
like K mean clustering and algorithm start analyzing the data on its own to form customer groups. For example, it
might figure out okay, so these are the customer who spend a lot and shop often. they are your high value buyers or these
are the one who only buy when there are discount and they're budget conscious shoppers and maybe there's a third group
who order just once those are your impulse buyers maybe you can create more custom offers and target these impulse
buyer to buy your product so basically in unsupervised learning you didn't tell the model what kind of customers you
have it discovered them by drawing insight from their behavior and grouping them together and once you have these
insights you can make smarter decision you can show personalized ad recommend better product and create offers that
actually match each group's buying style. This is the simple intuition behind clustering algorithm. You need to
learn different versions of algorithms such as K mean clustering, hierarchical clustering, DB scan clustering etc.
Reinforcement learning. Now think of a machine learning through trial and error like a game. The model which is agent
takes action get rewards or penalties and learns the best strategy over the time. This is no fixed data set. It
learns from the experience. This is how AI learns to play games, drive cars or manage stock portfolios. Moving on to
federated learning. Lastly, federated learning help machine learn without sharing your data. Instead of collecting
everything on one server, the model trains directly on devices like your phone. Only the model updates are shared
keeping your data private. It's widely used in app like mobile keywords or health tac. Common tools include
tensorflow. To start practicing, try these tools. Scikitlearn one of the best libraries for beginners. simple, well
doumented and packed with all the essential ML algorithm for classification, regression, clustering
and more. KAS, a highle deep learning library that is beginnerfriendly and built on top of TensorFlow. Perfect for
building and training neural network with just a few lines of code. You can learn ML from Google's AI machine
learning crash course neural network zero to her by Kapati. You can learn from Intellipar's YouTube video. Now
coming to step four, understand artificial neural network and dive into deep learning where you will have to
learn about CNN and RNN. Now let's step into the real brain of AI which is artificial neural network. These are the
foundation behind many gen AI application like chat GPT and more. Let's break this down with a simple
example. Cat versus dog image detection. Suppose you want to build an AI that can identify whether an image is of cat or a
dog. You start by feeding the model thousands of labeled images of cat and dogs. These images enter the input layer
of the neural network where each image get converted into grid of pixel values which are numbers. As the data moves
through multiple hidden layer, each layer tries to learn something from the image. One layer might detect edges,
another might identify ears, tails or fur patterns. The deeper you go, the more complex the feature becomes.
Finally, the output layer gives the result. For example, predicting whether the image is of a dog or a cat. Now, if
the prediction is incorrect, the network doesn't stop here. It learns from its mistake using a technique called back
propagation. We have a complete video on back propagation. You can check it out if you want to learn about the same.
Where the model calculates the error and adjust internal connection which is called waves to do better next time. The
math behind this adjustment is called gradient descent. It helps the network make tiny precise improvement to reduce
the error. By now you understand what a basic neural network is. But when it comes to genai, especially for working
with images or text, you will need to dive into two powerful types of network which is CNN and RNN. They power tools
like chart GPT and live translation apps. CNN or convolutional neural networks are great for image task. CNN
are used in face recognition a RNN or recurrent neural network work best with sequence like text or speech. Say your
model is completing a sentence. It needs to remember earlier words to predict the next. RNNs have memory for that and
better versions like LSTM and GRU help them remember even longer. They are used in chatbots, translation tool and speech
recognition. You can start with a project that predict the next word in a sentence using PyTorch or Keras.
Understand autoenccoders and transformers. Now we dive into architecture that changed everything
which is transformers. This is the model used in all major geni tools like GPT, Claude and Gemini. Transformers
introduce the self attention mechanism which helps the model focus on the most important part of the input. Before this
model struggled with the long sequence transformer fix that. Now what's an LLM? It stands for large language model. It's
basically a massive transformer trained on tons of text data from the internet. LMS can write poems, answer question,
explain codes and more. Understanding how they work from tokenization to embeddings to attention layers give you
real power as GI engineer. Key concepts in transformers include tokenization breaking input text into smaller parts
like word or subword embeddings turning those tokens into vector or numbers that model can process. Then self attention
the magic behind how model focuses on important word. So transformer are trained on massive data set using huge
computational power and output is what we call as LLM. Step five dive into generative models. Generative models are
what make genai different. Instead of just classifying data these models create new content. You will learn about
G or generative adversial network where two models are generator and discriminator compete. One tries to
create fake data and other tries to catch it. This back and forth make the generator smarter. There are also other
types of VAEEs and diffusion model. These models are used in AI art, defix, fashion and more. If you want to work in
creative AI, this is where your journey begins. Step six, learn prompt engineering. Even without training
model, you can get amazing results by mastering how to write prompts. Prompt engineering is like giving precise
instruction to your AI assistant. It's not just about asking question. It's about guiding the model step by step.
You will learn to use context, example, tone, and chaining techniques. The skill is super valuable if you're building
tool that rely on LLM. It's also helpful when you're working with APIs from open AI or go ahead where the right prompt
means a difference between a good and a bad result. Step seven, learn to use APIs. So most company won't just ask you
build GPT from scratch. Instead they will ask you to use the existing APIs. That's where this steps comes in. You
will learn how to call open AAIS GPT, Gemini by Google or cloud via their APIs. You will build web apps or tools
that use these APIs in background. For example, a chatbot, a rum writer or a meme captioner. Use programming
languages like JavaScript or Python for back end and front end. So once you learn how to send a request and get a
response from the model, you can build real product. Moving on to step eight, which is fine-tune elements. Fine-tuning
means taking an existing model like GPD2 or LMA and training it on your custom data set. Let's say you want a chatbot
for legal advice. You feed it case file, legal terms and previous judgments. The model learn from the data and become
specialized. You will use tools like hugging face, transformer, lora and pft to fine-tune efficiently. This step lets
you build highly customized AI tool that work for specific industry or user. Moving to our step nine, which is
explore multimodal AI. The future is not just about text. It's about combining text, images, audio, and video. That's
what multimodal AI is. Imagine uploading a photo and having the AI write a story about it. Or you speak a command and the
AI draws for you. Model like Sora, Gemini, and Dali are already doing this. The exciting part, you don't have to
build everything from scratch. Platform like hugging face and tools like Langchin lets you existing model
fine-tune them for your own needs and build custom AI that solve real world problem like automating customer
support, content generation or even healthcare chatbot. Hugging face is like a huge library of pre-trained AI model
for text, images, speech and more. You can simply become a model, test it online and plug into your project
without heavy coding. Lang chain on the other hand help you build AI app that can reason, take action and use tool
almost like a brain for your AI system. It connects model with memory, API, search tools and lets you design full
workflow with multiple steps. This is also where agentic AI comes in. AI that doesn't just answer question but can
take action. Think of creating your own AI system that read emails, searches the web, book appointment and even talk to
other tools all by itself. You will need to learn how to combine different input types and build tools that can handle
them. This opens up creative possibilities that go way beyond traditional apps. Now coming to our
final step which is step number 10. Now that you understand the basics of Genai, it's time to build real project. project
show what you can actually do. They are the best way to prove your skills. Start simple. Build a news summarizer using
OpenAI's API that turns long articles into a threeline summaries or a rum rewriter that takes a job role and
rewrites your resume using GPT. Want to try something visual? Create an image generator using Dale or a widget
classifier using TensorFlow and Steam. You can also build fun stories like AI story generator that writes story from
topics or a chrome extension that rewrites emails using GPD. All you need is basic Python APIs like OpenAI or
Hugging Face and simple tool like Flask, Streamlit or Gradio to bring your ideas to life. Start by breaking problem into
step input, model, output and build each part one by one. Now let's talk about hot project areas. One of the trending
ideas is MCP or multimodal conversational platform. These are AI tools that understand text, voice and
images together. For example, you speak a prompt and the AI replies with a story or image. Use tools like GPT DALI and
connect them using lang chain. You can also try projects like medical Q&A, B train on health data or a voice to image
generator that turns your spoken word into picture. Once you build something, upload it on GitHub, deploy it using
gradual or hugging face spaces and make a small demo video. This helps recruiter or client see your work in action. With
this we come to the end of the video and all the steps mentioned above are explained in detail in Intellipath
generative AI video which is available for free. You can watch it using the link provided below. Plus you can get
the complete road map in the description below. Just a quick info guys, Intellipad offers generative AI
certification course in collaboration with IHub IIT RUI. This course is specially designed for AI enthusiasts
who want to prepare and excel in the field of generative AI. Through this course, you will master geni skills like
foundation model, large language models, transformers, prompt engineering, diffusion models, and much more from top
industry experts. With this course, we have already helped thousands of professional and successful career
transition. You can check out their testimonials on our achievers channel whose link is given in the description
below. Without a doubt, this course can set your careers to new height. So visit the course page link given below in the
description and take a first step to a career growth in the field of generative AI. So now that you guys know the road
map to become a geni engineer, it's time that we get started with mastering the right tools and concept. For this I will
be handling over the next section to an industry expert. He will walk you through the essential from an
introduction to generative AI and transformers to open AI's GPT lchain and craft prompt engineering. So let's get
started. >> So uh we'll be covering the following topics. Um as far as uh
um you know this you know my course is concerned what are the topics that we're going to be covering? Um we will be of
course be covering uh I I'll start with uh an introduction
to generative AI right so we'll be doing that in our today's session um we'll be
very high level broad brush strokes um we'll be discussing introduction to generative AI in our today's session um
and then what we will als also be doing is we will also be covering topics specifically around
um why do I see that folks are saying there's an eco so introduction to genai
so I'm going to be talking about all the different applications of u genai right so how the industry is perceiving
uh an industry point of view so we'll be covering these topics um broadly in our today's session um more of a business
point of view right so where is this this area where is this field sort of headed towards and stuff like that
that's what I'm going to be covering broadly in our today's session then from after the session we'll be going into a
lot of detail right so I will talk about um uh the transformer architecture right I'll be talking about transformer
architectures um I'll also be talking about how some of the most popular GPD models
are trained right um and then we will also be discussing uh we'll also architectures are um you know how they
work then we will go one level lower um we will actually start discussing about um you know uh we'll be doing a lot of
hands-on specifically on trying out some of these architectures. So I'll introduce you to firstly the the open AI
uh so we'll be I'll be focusing on the open AI models uh for for for a good chunk of this particular course. Um I'll
also see if I can show you some open-source models, right? So there are different types of u different ways you
can access some of these models. So um I'll also I'll focus primarily on the open AI model but also show you how you
can access the u other models that are available out there. So the open AAI GPT models um um is you know how to access
them. Then uh I'll introduce you to lang chain uh which is a library that is very very popular. It's an orchestration
library that helps you access some of these models uh very efficiently. Um then we will look at
um you know u some prompting techniques right so I I I'll prompt engineering to be specific right I'll I'll discuss
about uh some topics around prompt engineering all the different prompting techniques that you would typically have
so chain of thought right um I I'll talk about react um we'll also talk about tree of thought
um and so on and so forth. There's there there's a bunch of other things. So we'll talk about all of those uh few
short learning and stuff like that SSL and stuff like that. So we'll talk about all of that u under prompt engineering
and then once that is done we will then get into retrieval augmented generation
which is also popularly referred to as rag. So we will talk about rag. we will understand how rag works and then we'll
do a lot of hands-on on rag as well. Um and then after rag um I will also show you some more complex um u you know
agentic architectures or rather simply put let's say agents um using langra
um and stuff like that. So, so I'll probably be closing out the sessions at the end using agents um and and concepts
of agents in Langra. So, broadly this is how we will go about doing things. Um in the later parts of the session or maybe
actually here itself when I discuss OpenAI, I will discuss of course the GBD models, I'll also discuss some of the
image generation models here as well. Um so how you could access the dolly kind of models how can you actually um
generate content using the dolly kind of models I'll also be discussing that um in that session so so broadly these are
the topics that I'm going to be covering um again I don't want to while we are covering one of these topics of course
we'll end up covering some of the ancillary topics as well right so topics around this space as well I know this
might not be I mean all of these topics that you're currently seeing on the may or may not necessarily resonate with a
lot of you because you may not know what this space is but but trust me this pretty much covers the 80% of uh I would
say everything that's out there today right so a good 80 is you know 75 80% of all the happenings in this particular
space is fairly covered in the topics that you see on the screen over here let's start with the introduction to
Genai so so all of this by the way is LLM's only so when I say open AAI GPT GPT models. These are large language
models only. When I'm talking about transformers, those are large language models only. So
I'll talk about all of that as we speak. So here's the thing, right? So so let's let's start with the first topic today,
right? Let's talk about what is generative AI? Why all this drama about geni? Why has it suddenly become so
popular? Right? Let me agenda introduction to Genai. Perfect. So while I bring this up, I also want to bring up
some presentations. Give me a second. So one of the good advantages of being uh in the industry while you're doing this
is is that I also get a lot of content from uh a lot of these um
companies out there. Let me show you some interesting content that I had got very recently from Bane, from Microsoft,
from Accenture. Some lot of very interesting presentations out there. I'll try to
bring some of that up. Easy to understand kind of slides or easy to understand kind of content. So you all
have you know when we talk about jai or rather if you kind of talk about what has changed over the last couple of
months couple of maybe one and a half year or so one year or so I would say you've suddenly seen these tools that
you see on the screen suddenly pick up pace all of us agree on this right u yeah bar has become Gemini um chat GPT
has become so popular. Google has also launched Palm kind of models. Um, Anthropic has launched uh an interface
called Claude. Um, Perplexity is another tool. You have OpenAI has launched Dali. Um, Microsoft has launched Copilots. Um,
yeah. The point is we suddenly saw something change in the space, right? So cut to two years, 3 years ago, we are
still we were still talking about all right, how can I build a a deep learning model that can do
question answering, right? Or how can I build a recurrent neural network based model that can do classification?
How can I use some of the existing B models for doing sentence similarity or document similarity? How can I use word
toe to let's say do something very specific in this particular space like maybe
classifications or let's say document uh similarity question answering so on and so forth this is what we were talking
about two years ago but suddenly things have changed suddenly you start you know Chad GPT was launched and and we have to
admit that Chad GPT's launch is a marquee event in the history of AI so far right chat GPT launching is
something that will get etched in the books of history, the AI history forever and ever, right? So the moment Chad GPT
got launched, it just, you know, it was a jaw-dropping moment for everyone. People suddenly started seeing some and
I I I still remember a lot of these uh posts coming up on LinkedIn, on Twitter and everywhere where people started
saying this is revolutionary. this is going to change how we look at AI anymore. This is going to completely put
people out of jobs blah blah blah. That fear is still there even today. So this is what a outsider is looking at, right?
So you're you're using the chat interface through chat GPT. You're asking it questions. It's able to do an
amazing job of answering some of these questions, right? It is able to create content. It is able to surprisingly
create content with amazing levels of accuracy. um levels of accuracy that are far
beyond what let's say uh any human or uh uh decent it would very easily build a you
know beat a decently skilled human in some of these some of some very very specific tasks you had seen these genai
models beat um people in or rather students in bar exams and SAT exams right? Um stuff like
that. Question is like where did this suddenly come from? Like was this something that happened all of a sudden?
Does this have anything to do with anything that we learned so far or is this completely new? Right. Um the
answer to it is a bit of both. It it has a lot to do with what we have learned so far yet it is completely new. Um it has
it has lot of eerie similarities to stuff that we have discussed that we may have you know spoken about so far yet
this has um you know structurally fundamentally you know even conceptually this is very very novel very new in this
particular area. um and that is why you suddenly saw this spike in the number of apps that started
coming up, number of know chat interfaces that started picking up and so on and so forth. So long story short,
new new field of AI has suddenly emerged u new tools came in and suddenly it seems like AI has become a lot more
easier than what you and I would have imagined so far. it suddenly feels like man why did we learn all that we've
learned so far right this seems super easy right you could just you could just open a chbpt interface and get it to do
things which you would have otherwise struggled quite a bit uh with with traditional AI system so suddenly even I
went through this brief period saying okay am I going to go out of job because I I I spent a lot of time in AI and am I
going to go out of job and I decided otherwise to to kind ride the wave rather than try and sit
down and sob about it. That's a different story. We'll come to that later. But here is where we are. Okay.
So now the question is what is it about genai? What is it? How is generative AI similar or new in whichever form or
shape, right? So let's talk about that. Um so you if you remember we spoke about this prophecy, right? So
this AI prophecy or this AI um ambition, aspiration of being able to build something like a
Jarvis, right? Something like Terminator, artificial general intelligence. You all
must have heard about artificial general intelligence. So what is artificial general intelligence? So when you talk
about AI, when you talk about artificial intelligence, AI can be nicely split into two parts. artificial
narrow intelligence and artificial general intelligence.
The key operative difference between the two being here artificial narrow intelligence and artificial general
intelligence. What are what is the difference between the two? See artificial
the difference between artificial narrow intelligence versus artificial general intelligence
um is very simple. So, so far we all have spoken about all of these different
applications of you know of AI like whatever you see on the screen right be it let's say your portana your Apple
Siri your Google now your even your self-driving car recommendation systems on YouTube your um AI capabilities on
you know your machine learning AI capabil ities on your Uber app, any kind of AI that we may have done so far, ADAS
as well. Adas is a very very good example as well, even ADAS uh which has LAR um you know, you know, ultrasound
capabilities to identify what things are around it, stuff like that. So all of these are great AI capabilities, but
they are all artificial narrow intelligence. What do I mean by that? each of these activities. So take for
example YouTube as a good example. If you take YouTube, YouTube has a lot of intelligence in one very good example is
is recommendation engine. So the recommendation engine that you have that is recommending the next video that you
should watch on YouTube is only built to do just that piece of work. It can only do that part of it. it cannot I cannot
use that same piece of AI or that same model to do let's say to identify what maybe um I cannot use that same AI to to
get to get to to predict how long will you watch that particular video for or I cannot use the same model to identify
uh if you would purchase let's say the YouTube premium subscription or not, right? I cannot let it do multiple
things. It can only do one very very specific activity. It'll do a fantastic job of that one activity, but it cannot
do anything beyond that. It is just trained to do that one piece of work and it'll just do that one. It cannot help
you. That AI model which uh which predicts um if which video you will likely to watch cannot tell me cannot I
cannot use that to generate subtitles. I cannot use that to summarize that complete video. I cannot use that same
model to do other things. You can only do one piece of work. To do the others, you'll have to build other models.
You'll have to build separate solutions for it. And that's how we've been doing it so far. To solve a specific AI
capability, you need to build one model for it. You cannot have one model that does everything or one model that has a
wider application. It is very narrow. It's not to say that it is bad, right? I don't want you all to misunderstand that
this is bad. This is not bad, not at all bad. This is great. It's just that it is capable of doing one particular task at
a time. But what we are but the but the objective of the field of AI is to try and get to build something like a
terminator artificial general intelligence. I want to build or rather we intend to build I don't know if you
watched Dennis the Menace. I want to probably build a robot that has got that kind of intelligence in Dennis the
Menace or maybe in Jetsons for that matter. You we've all we've all grown up seeing that kind of intelligence. Right?
So in in in science fiction and movies some form or the other Skynetss, Terminators, your Jetsons being another,
you know, smaller silly example. The point is you want to kind of get there. Uh small wonder, not little wonder,
small wonder. >> Yes. Um Vicki from Small Wonder, one of my favorite shows right at that time. Um
all right. So point is how do you get there? So today we don't have any applications of artificial general
intelligence. Right. Um we are trying to get there. We are far away from it. we are a good few maybe I would I would
still assume like a good few decades away from it but that's where we want to get to that's what we would want to
really try and aim at right artificial general intelligence what you and I have is AGI right um we can do multiple
things all at the same time I can drive a car I can cook food I can teach my child I can I can learn something and I
can also attend a session in parallel I can do so many things in parallel with decent amount of accuracy across
everything right um that's what we're trying to get at artificial general intelligence it's very humanlike
intelligence now the thing about um AI my friends is that uh AI is such a new space
um you know when you talk about um you know when you talk about artificial
intelligence um we are more often than not always talking about artificial articial narrow
intelligence capability. We're not talking about artificial general intelligence capabilities at all, right?
Because we're a couple of decades away. But what generative AI has done is it has gotten us, it has helped us make a
huge leap towards AGI, but we're still far away. But it has gotten us far more closer to
you know to general intellig to general intelligence than what we would have done had we progressed in the current
phase. If we were to build the same kinds of pieces of technology like what we're currently doing it would have
taken us a very very very long time to get there. But at least we were able to make a massive stride in that direction.
Why? What is it about generative AI which makes us believe that we've gotten closer to general intelligence? Remember
everyone, I am not saying we've made progress on AGI. We haven't. We are still talking about narrow intelligence,
but we've been able to make much larger progress towards AGI using generative AI models. I'm not saying by any means that
we've gotten closer or rather we have made progress on AGI itself. The objective of a company like OpenAI
is to be able to build AGI. That's what that's what S Alman has always talked about. He always keeps saying we want to
build AGI. We want to build artificial general intelligence. But I not just I mean who am I, right? And I'm a small
fish in the pond. But if you take for example the the who's who of the industry out there, they all believe
that we're far away from it, right? We're at least a good decade decade and a half away from it. we need more such
wonders like Chad GPT to happen before we can get there. So that's essentially
um if you were to consider let's say colonizing um let's say another planet as as the
end objective our Chandrean 3 was a huge leap towards that consider Chad GPT or the generative AI as your Chundrean 3
success u have we gotten closer to colonizing the planet maybe yes but have we accomplished anything in that
objective probably not to make a huge we a huge leap towards that direction. So that's a fair analogy as far as how you
should treat generative AI. But the beauty of this space friends is that with that small leap itself, we are
seeing such huge or huge changes in how the industry has started to use AI in the regular in their day-to-day basis.
The good part is you and I we have the object, we have the um opportunity to ride the wave and to kind of stay on top
of it, right? So we you are all kind of getting into the field of AI exactly at the time where this transition is
happening. We've kind of made that leaprog. It's that it's that leaprog moment. What is it about um generative
AI that makes us believe that we've made that? Let me explain. Let me show you a couple of things um why why we say that
um there are certain very interesting capabilities of generative AI that makes us believe something like this. Let me
show you an interesting presentation from Microsoft. This is a Microsoft presentation. Uh they had actually
presented this to us uh in in it's a public presentation. So on what is it about Genai
that uh makes us believe um that it does things differently.
see generative AI right and these are examples right um it can do
far more things than what we are currently seeing on the screen of course um but for example you all have used the
chat interfaces right so this is a very very good example of it so if you take for example
the prompts um so you can actually just ask it a question uh and it can simply respond back to
You can follow up on those questions as well. You can sort of have a a a genuine humanlike
chat and it'll actually respond back very very much like a human being. Right? So its understanding of language
has significantly improved since the last time. Right? So models now understand language very very very well.
Right? than they would have done probably using any of the other regular AI models. Right? That's number one. So
understanding as language has significantly changed since the last u since the time we know what uh has
happened in this particular space. Um or even if you take for example the um your BERT models, you take any of your
other models that you've done learned that you may have learned about so far, your these models are far better at
understanding language. Not just that, these models are also very very good at understanding code, right? They're not
just good at understanding language, they also understand code very very well,
right? So they can actually write code for you. I have I cannot tell you how many times
right so I'll tell you a good example so today we built a capability in our in my own team we built a mobile we built an
app um in my in my in my team a product in my team and then we've decided that we should do a
complete code rehaul so we've written a lot of stuff using object-oriented programming and we decided you know what
oops in our setup makes no sense let's change it to functional programming and we just took that complete code uh bit
by bit we put it on chat GPT asked it to convert this into functional code 20,000 lines of code was rewritten
by shat GPT in a mere 3 days 20,000 lines of code we were able to just restructure our
complete code base in two to three days it's just amazing just the kind of capabilities that we starting to see um
that you know that that uh u something like a chat GPT could do. Those are the kinds of things what I you know I al I
can also talk about how content generation has changed right so one of the things about generative AI because
it is able to generate content. It is able to create new content. We can also get it to create images. I can tell it
what image I want and it actually create that beautiful image for me bases what I asked it to create. Um and and and and
not just this now you also have videos that you can create as well right so you can also get it to create some videos I
can also use Sora and stuff like that that I can also create too as well I I'll show examples um I'll show a lot of
examples as we go along for the next couple of weeks I'll just be showing a lot of hand hands-on examples itself so
so don't worry again I want you to understand that these are the kinds of some of the applications of of uh of
generative of AI. Uh when I say content creation by API, um API essentially is simply nothing but an endpoint. Um you
could have a model that's sitting somewhere and um you can interact with it. Don't worry about this this line
over here. I'll exactly explain what I mean by it. Now I want to show you something slightly more interesting. So
let me let me pull up um another interesting presentation. This is another presentation by by McKini. Uh
again very I want to talk about a couple of slides on this. Um lot of companies these days right um they have been using
traditional AI right so they've been doing pattern recognition um and and and stuff like that like
traditional AI capabilities. Now with GI you could do a lot of you could do code generation you could do image generation
you could do enhanced pattern recognition some of the capabilities that you we would have probably wanted
to have solved earlier we would have to solve it much le you know we can solve it much faster right now why are these
models so good right so why are these models that we speak about as far as uh geni so so damn good right again
approximately What you see here is and this is again a public presentation by me. So you can
actually download it from the website. So what they're saying is u these models are so good because they've been trained
on massive volumes of data right so for example one example here is if you take any of the uh GPT models GPT3 for
example was trained on 45 terabytes of data right uh a complete crawl of the internet complete internet was used uh
some lot of Reddit uh content, more than 250,000 books, the whole of Wikipedia, man. That's that's almost all of the
internet that you and I have access to. All of that data has been taken and they have provided that to the model. The
model has a very interesting way of learning. Um we'll talk about how these models learn. Um the way they learn is
very much to predict the next word. Right? So given a particular sentence, given a particular word, they try to
predict the next word. So when I say the models have been trained, they've been trained for a classification problem to
predict what the next word is. Given a set of set of sentence, given the first word, predict the next word. First two
words, predict the third word. First three words, predict the fourth word and so on and so forth. So you're sort of
predicting always the next word in the sentence. So point is that model is 175 billion parameters. When say 175 billion
parameters, what do I mean by parameters here? Weights and biases. These are your weights and biases. A total of 175
billion weights and biases not million billion. Okay, not features. They are not features. They are the weight and
biases. They are your parameters inside the model. Right? And had you trained this 45 billion terabytes of data with
75 175 billion parameters on one GPU, it would have taken you like 32 years to train that model. It would take you 32
years to train that model, right? So you can imagine if you were to train it for 32 years, then you would never get the
model. So what do you do? How do you accelerate that training? You either reduce the data or reduce the model or
throw more compute at it. Right? I'm I'm I'm kind of oversimplifying it here, but when I say throw more compute at it,
basically get get more GPUs. Now, it's that part which is where a lot of AI wars are happening right now, right? Of
course, to train such large volumes of data, such large models, you need more and more GPUs. But to get more GPUs, you
need um GPUs are not are are not made randomly, right? So, GPUs are very expensive. You need to make them. So,
who makes these GPUs? It's Nvidia's of the world. It's the AMDs of the world. They are the ones that are making these.
And Nvidia has is in such a beautiful spot here that they have uh really catched the cow left, right, and center,
right? So, they're doing a very good job with how they're positioning themselves. Anyway, we we'll talk about that later.
Point is again that it's because of that reason because there is a need for compute you started to see that you need
more and more models more and more compute um sorry more and more GPUs to accelerate the training but then even
after that right so let's say you do all the training what do you get you get a model which is 800 GB in size very
simple GPD3 is 800 GB can you imagine a model that's of the size of 800 GB it's that big it's not storing data data is
45 terabytes the the model is 800GB. So, so it's just these 175 billion weights and biases that are being stored um and
and and the size of that is approximately around 800 GB. Can you load an 800 GB model on your machine on
your personal machine? Of course no, you just cannot, right? Your machine has 16 GB RAM or 32 gig memory. Maybe if you're
rich enough, you'll probably buy a 128 gig memory machine. How can you load a 800 GB machine? How GB model? That is
where API is coming. before. So the model cannot sit on my machine. I cannot download the model unlike how you and I
were building models earlier where we took the data, we loaded the data in our local, we put it in a folder, we take
the we take the code and we actually build it uh in our machines. Um you install the libraries on your machine
and you build the model. There is no more model building, my friends. Model is not even going to be
built by you and me anymore. model is built by Microsoft, by Amazon's by Google's by Metas. They are the ones
that are going to be building the models. You and I are not building the models anymore. You'll be using the
models that they've built. Model building, not so much any of our job anymore. When I say our job, of course,
unless you choose to work in that space where you want to build models, that's that's a different story. My point is as
end users you and I are not building the models. You and I are using these models.
Um but where are these models? These models are going to be hosted somewhere else. You know in cloud, right? You can
I mean you can load these models but you will have to set up a data center to host a 800GB model. You need a data
center. You need a huge data center that can that has 800 GB in terms of RAM. you need a you need racks and racks of RAMs
to um to to support an 800GB model. What do you do otherwise? You put it on cloud. You let you let Microsoft manage
this and say, you know what, I don't care how you manage it. This is the model. You manage it. OpenAI has said
Microsoft, you manage this model for me. You host this model. Provide an interface where I can access this model.
Right? I want to access this model over internet. So, I just just give me a layer just like how I access a website.
I want an interface. If I want an API layer, that's where APIs come in where I can just simply query and I can get the
response. I don't want to be doing the business of actually loading this model by my my own self. So this is almost
like a service model as a service. Exactly. So what do you see here? See, geni my friend is not new. Jai is new is
not at all new. And this is the part that I was talking about, right? If you remember I spoke about, hey, is Jai
completely new? Well, maybe not. It's it's actually not completely new. Um, it is been there for a good seven, eight
years now. So that this is a good ex good view of all of that. The model the foundation of this particular model was
first published um in oops the first published in 2017 that to December 2017 a p paper called
attention is all you need is the paper that was first published. What happened after that? This paper took the world by
storm. Why? This paper is everything that kind of completely changed. Right? This is a game changer. What happened
here? So I again we'll I'll talk about this in much more detail when we discuss the transformer architectures. This
paper is the is the is the paper that introduced transformers to the world. So I'll give you one example. So far you
know um you all have learned RNN's I'm assuming right? You all have learned RNN's recurrent neural networks in an
RNN or LSTM for that matter or any of your models. So if you take for example a sentence which has W1, W2, W3, W4, W5.
So let's say five words or how many other words that you have in an RNN or an LSTM. You're essentially passing the
current word as an input, right? Um and then you could potentially predict the next word as an output, right? And then
of course you have a you have a bit of recurrence over here which is let's say an RNN or an NSTM so to say. You could
pass the current word as an input and you can predict the next word as the output. Then you can take these two
words as the input and then you can predict this one as the output. Then you can take these three words as the input
and you can predict this one as the output and so on and so forth. So the training in this particular case, right?
How did we learn the dependency between one word to another word? You of course had to go from you had to take the first
word, predict the next word, take the next two word, predict the next word, take the next three words, predict the
next word and so on and so forth. So you had to learn it like a language where even you had to move from left to right
or whichever way left to right, right to left, whatever. Point is that you had to sequentially learn the data. Language is
a sequence is a collection of words. There is an element of sequence associated with it. So you kind of learn
from left to right. Now in the process of learning from left to right especially when you have large volumes
of data there are issues of losing dependencies right when you have that's why LSTMs
came in or DRUS came in to kind of address for those long-term dependencies and so on and so forth but you still
were learning sequentially so when you learn sequentially it was very slow the learning process super slow but you
still were able to do a decent job of the whole learning process what but what these new models did firstly What
transformers did differently, right? Is they completely took away this concept of learning sequentially. They said no
more learning sequentially. You don't need to learn sequentially anymore. Right? Why do they say that? They say
look given any particular word this word has dependencies before this and after this. So that's where they introduced a
concept called as attention. Attention is a new concept. What it tried to do is it tried to say look when
you look at a sentence when you read a sentence if you take any one particular word in a sentence. So for example if I
say I had an amazing day at the park rather an amazing um right I had an amazing day at the park.
If you take a sentence like this every word in some form or the other has some kind of a dependency on the other words
that you see over here. When I say I, you know, I is of course partially dependent on some of the other words. If
you take the word amazing, amazing is is talking about the word day. Amazing is talking about the word I. Amazing is
also somehow talking about is also addressing the park, right? And if you take for example park, park is again
dependent on the word amazing. Park is dependent on somehow the word day and so on and so forth. Point is when you take
any sentence every word in a particular sentence has some form or the other some kind of a dependency with its
surrounding words. So and and those dependencies if you are able to capture it differently
rather than simply just trying to learn sequentially. If you capture those dependencies differently and you capture
all of those dependencies to some kind of a numeric representation like some kind of a more efficient embeddings.
What you could possibly do is you can just take these embeddings and then simply pass it into the model. So you're
eliminating, we'll discuss this in much more greater detail when we discuss transformers. But the point is you're
eliminating this sequential learning aspect of models. When you eliminate the whole process of sequentially learning,
you are accelerating learning. Um and and the concept of this exactly you can learn parallelly. And when you can do
parallel learning, you can do you can do a lot more epochs. Um your finetuning or rather your tuning could be much faster.
the the complete math becomes so much more simpler. One of the problems with LSTMs and RNN is because the math has
become so complex that you know you cannot learn with large volumes of data. But what transformers were able to do is
they kind of broke that complete sequential learning down into a very simple uh you know simple learning
process like how you would learn any of the other uh convolutional neural networks or even for that matter a
regular feed forward neural network. They just made the whole learning language a very simple exercise. Because
of that, these models have become super I mean a you were able to learn language very fast and secondly um you could go
into far more um you were able to also extract a lot more detail about these transformers as
well. uh we'll talk about transformers in much much greater detail but the point is the point that I want to make
here is these transformer models because of the introduction of these transformer models it completely revolutionized the
whole learning process of language learning language became so much more faster so much more simpler um and that
made it that that kind of added so much more uh ease with which how people could learn or how these models could learn.
That was one of the reasons why the 2017 paper is a massive massive leap. It's almost like somebody was saying, right?
It's it's almost like I I was reading about it somewhere and somebody said it's almost like somebody from the
future kind of came because this architecture is so different from how we have learned neural networks so
far, right? We it's it's very very radical. It's very different from how we have learned language so far because all
that we've learned is we've learned RNNs, we've learned LSTMs, um we've also learned encoder, decoder,
sequencetose sequence model. Point is they're all learning sequences and suddenly in 2017 somebody publishes this
paper and says all that you need is attention. Forget about sequences. Sequences is doesn't matter. You just
need to know how to encode those dependencies very well at the beginning itself. If you encode those
dependencies, everything is addressed. Which is why the paper says attention is all you need. All that you need to do is
some time of attention, everything else can be taken care of. So this paper was so radical it and somebody was also
saying right I was listening to a podcast. It was almost like somebody from the future came in and kind of
whispered this architecture in somebody's ears and then left. I I'm currently watching uh Lord of the Rings
uh the the the Amazon show uh that that u rings of power um it's almost like how the rings were you know formed right so
you have Sauron coming in and whispering it in the ears of calibb on how to uh you know forge the rings exactly like
that it's almost like that like somebody just came in and whispered in the ears of Google and said is the is the fallen
angel right exactly so it's almost uh somebody just came in and whispered how how this architecture sought to be
created. It's brilliant the way they've written the architecture and the point is that once that's sort of created
um once that architecture was published you could see what followed after that six years down the line my friend you
have seen one of the most revolutionary products that came out from there on right so you have um you have of course
the um open AI models coming up you had uh you know llamas coming out by meta uh um
uh Hugging Face launched their models, Anthropic launched their models and so on and so forth. So many models got
launched, so many models got launched. This was still very old. Sora came in um lot of things happened in this
particular space. So yeah, I mean so many different people kind of building their own models right now. uh even in
you know as we speak today um and this was uh you know I think one of the recent uh posts by uh by Eli on uh how
AI is evolving in India. I can actually show you that as well. Let me just Yeah, this is the India outlook for uh Genai
by Ernston Young. uh they very recently published this um fairly very cool uh on how they kind of published it um they're
talking about how um you know some specific areas where I think uh let me just go into this a little yeah so a lot
of people focusing in Indian organizations a lot of the work that's happening right now coding assistance
document intelligence and so on and so forth a lot of interesting work that's happening already there are companies
that are building these models itself. Somebody spoke about uh Bah GPT. Yeah, there's a lot of lot of companies
building their own um AI models as well, right? So um again, super super fast. By the way, these these images are all
created by OpenAI. Um all the images that you're seeing here, they're all created by them. Here are some um Indic
LMS right so Amari Canada Openhati Bhashini this is the the government of India's uh capability called Bhashini
um yeah so many so lot a lot of Bharat GPT somebody spoke about Bharaj GPT um yeah here you go these are all the
different uh capabilities in the India space as we speak today um Sam is the was actually the first company called
openhati um um they actually did a fairly good job as well and all of these companies
are u doing a very very good job right now and in my opinion they're doing a fantastic job as we speak of of uh of
AI. All right. So I I I don't want to go into too much uh of the detail around how AI is working but as you can see
there's a lot of Indian companies as well in this space right so that are doing it that that are doing a very good
job right so there's a lot of Indian ventures uh AI ventures that have been doing a fantastic job in this particular
area as well um they they've also secured a lot of funding by themselves but yeah okay uh let's let's go back to
what we were discussing so a huge advance adancements in the field of AI. That's what I wanted to talk about since
2017. Yeah, I don't want to get into this but I just want to talk about this this particular slide is again very
interesting in my opinion because see um the point that I wanted to speak about right who is building these models not
everyone right very few people are building these large language models these large gen models the only few
companies are building this and if you see those names at the top very few these are all big names right so these
are all big names that are being built. Of course, now you have also have these some of these Indian companies also
building their own models by themselves. But point is some models on text, some models on code, some on image, some on
audio, video, 3D, blah blah blah. So many different models that are now coming out. But all of these models in
some shape or form are all built on large large volumes of data. uh and these large volumes of data is exactly
what we are talking about as far as um or rather these these large models is uh um is is the size of these models is
exactly the reason why we require different kinds of uh infrastructure to handle. So the way you deal with these
models is going to be very different. The infrastructure requirements are different, the modeling requirements are
different. So if you're not going to build the model then how do you access these models? So what do you do with
these models? How can you customize these models for your own requirements? Because you're not building the model
and how do you get the model to work on your data? How do you get it to do what you specifically want it to do? Um so
those things we need to discuss and and there are some very new design patterns uh that have emerged over the years. Rag
being one of the most very very common very popular uh design pattern that has emerged in how you can how you can use
some of these models. uh but we can discuss them into much greater detail. Um so let's let's go back um to this. So
so what are we saying? We're saying look AI massive massive uh transformation in the
space. Um huge leap towards general intelligence. Lot of very interesting things that have
happened in the space since 2017. 2017 the paper art of attention is all you need was like the start of all of it and
then very interesting things followed. Now let me talk about the value chain of this. So who is making money the value
chain in this whole geni space right again a business perspective we'll get into the technical
details here but a business point of view very important for you to understand because again um this space
is super ripe. So who is who are the ones that are actually making a lot of money here.
So let's talk about different personas here right let's take couple of examples there are
uh you know sort of broadly speaking there are four to five different personas here or four to five different
types of stakeholders involved in this whole thing. At the most bottom at the bottommost level you have companies like
these are your um you know these are sort of your
cloud providers. So it can be companies like AWS, it can be companies like Azure uh or
Microsoft for that matter and so on and so forth. So AWS, Azure cloud providers, these companies are providing compute,
right? They are the ones that are providing the data center, right? Why are they important? Because
they are the companies that are going to be providing the actual compute for these companies to host the models. So
it's it's these cloud providers where the models are actually hosted or trained at.
But who is actually providing them these models? So who's actually helping them build these models? So of
course on top of this they are the companies or these research companies that are building these models. the AI
research companies like your open AI, your metas and so on and so forth. Of
course, Microsoft and they they all they also of course have their own um um they they of course have their own uh
AI research teams as well. Uh my point is it's these research companies that we are talking about over here. So um so
you have the AI research companies, you have the cloud providers, but under the cloud providers, there is one specific
uh you know another one specific piece that I'd want to call out here which is the hardware provider,
the OEM manufacturers, the hardware manufacturers, Nvidia. Exactly. Amir and and and and Ali, right? So it's it's
companies like Nvidia, the hardware providers. So you have the hardware providers like your Intel,
Nvidia, AMD and so on and so forth. Um cloud providers like your Azure, AWS,
right? GCP, research companies like Meta, Open AI, right? Uh
um you have of course Microsoft has their own research company research research division. Amazon has their own
uh research division and so on and so forth. Of course all of them Google. Yeah. So of course Google
uh Google brain and so on. Deep mind which is their research division so on and so forth. So again these are three
personas very important persona. Clear so far everyone? uh these are the three three types of stakeholders in this.
Then who else do we have here? So the models are built here right? So necessary computers provided models are
built. Now what do you do with these models then comes again two cohorts I would say from here
two groups from here right on one side these are companies that are not consulting companies these are
application uh or rather these are the product companies. These are AI product firms.
They are all the AI product companies and then you have developer tools or capabilities, right? So AI development
tools, development capabilities. What do I mean by this? So on one side you have you know people building your co-pilots,
right? your chat GPTs. You have um uh you have tools like for example fig
your Adob Firefly um Gemini exactly all of these you know chat interfaces your end consumer
products are being built consumer B2B B2C products that are being built that is all available here for sure
right um and on the developer side on the dev tools you're talking about um your lang chain and this is majorly open
source lang chain um your llama index and so on and so forth. So basically
these are open-source developer tools again because they are essentially going to be building them to build better
capabilities themselves right so they are open source models open source companies they're building platforms um
some are closed source platforms um like your lang so you want to build some AI models or whatever you want to
do some monitoring you would want to do some Um so you want to let's say host the
model in your environment you want a security around it um you want application level access controls and
stuff like all of those basically are going to be sitting here um and then on top of it u actually
these companies might also be using them but I just want to put them here because there is like sort of an overlap between
um because they are also directly working hard. These are also working here. Uh but these are also of course
using developer tools and they're contributing back as well. So it's almost like a producer consumer proumer
is is sort of a new word over here. But point is they're also interacting with each other. Then comes the topmost layer
and again over here they're probably again two boxes here, right? So B2B users or your uh final users
um and then there are also at times somebody spoke about you know consulting companies
and so on and so forth. So you can of course add one more layer over here in between which is the consulting layer um
and then you have the B2B users um as well that are sitting on top of it. So you might have users here or you can
also have users here depending upon how you're building these models. My point that the point that I'm trying to make
is these users are users like you and me. It could also be companies like your company, my company, my organization
where we are using GI to solve a specific problem and stuff like that. Um then
when you are talking about this value chain today so let's take a very simple example so imagine I have to let's say
build a small so let's say today one of the things that genai can do a very good job of is generate new content
right generate few new images for my marketing teams so my where are my marketing teams my marketing teams teams
are here. This is my marketing team that creates this content, right? So, how are they going to be using it? They
are probably going to be using they're probably going to be sitting here. They're probably going to be using one
of the tools that we have over here like Adob Firefly or something or maybe a chat GPT. Um, they would probably just
be simply using it. Those are built by uh either your Microsoft of the world or the Google of the world or Adobe's of
the world. They are built on top of your Azure or something or AWS SAS offering. Uh and they require
Nvidia, Intel, AMD. So today if you look at it, the companies that are making the most amount of money are here. It's
actually this probably is the only company that's making a lot of money. None of the others are making any money.
The cloud providers are not making any money. The cloud providers are in this are in this race
just to make sure they have a competitive advantage. Your Azure, AWS, they're not making any money out of this
today. They're charging you, but the reality is they are not making any money by them for themselves. They're probably
in fact even losing money in certain instances. It's these open AIs. Is open AI making
money? Probably not. Open AI is also probably not making any money. It's only the hardware companies that are probably
making money in this case. Some AI product companies may be making some money but very very early days right
some AI companies product companies might be making some money very early days and as far as the end business
users is concerned in specific areas yes the business users are making some money I would not say a lot of it um but they
are of course they are sort of making some money in specific areas they are making money for sure um again from
bottom to top the folks at the bottom will make always the most amount of money and the consulting companies are
make money anyways. These consulting companies will will make money on on your dead body as well. So these
consulting companies uh you know hopeless these baines and mckes of the world they'll make money anyhow.
Of course they're making a lot of money. Um they just have a new thing to sell. So then they'll sell that thing you know
to you. And it's not my perspective. The point is they of course will uh will make money any which ways because that's
their niche right how muchever we think they're useless they're equally useful because you need them to sell something
they have a certain brand value they you know they do sell a lot of things u you know especially in an enterprise setup
as much as I hate to have them they do a very good job of selling things within my company so I do use them in my own
company to sell things in my own company with my leadership especially um but the point is this is the value chain today
as far as generative VI is concerned. Uh it's who's making money is the hardware providers that are making money. Nvidia
is making money. Nobody else is making any money in this in this whole thing. Probably some of the end companies like
your company, my company is probably making a little bit of money, but otherwise u everybody's currently
spending money on this. Everybody's investing. Cloud providers are investing. They just they're not making
any money out of it today. Honestly, they just want to stay out stay out there. Uh but because if they've got a
foot in the door, they can make themselves indispensable. Trust me, they will um they'll make a
ton of money here. >> Maybe not today, but eventually they'll make a lot of money. The other thing, my
friends, is if you take a product like Chad GPT, right? You know, Chad GPT is built by
them. So in the case of Chad GPT, OpenAI is playing both of these roles. they're also giving you some developer tools as
well, right? So for people to to build some capabilities on top of GPT um so openAI has has a play in both of these
areas. If you take for example something like Microsoft copilot, Microsoft is playing the role all the way till there
right? So they have their models, they of course have their research teams that's building the product, they have
their product teams that's building the product and then they're selling it directly over here. So they're sort of
playing a much larger game. Nvidia interestingly also has a play here all the way up until the top almost here on
this side Nvidia has got some play. Uh but uh but that integration is not very strong but at least until the research
company level they do have a play. Um they are integrated until that particular point. Um so yeah Microsoft
by the way also has a consulting org over here. Microsoft also has a consulting or wherein they are also
making a lot of money all the way up until there. Um so their consulting team is also probably making some money out
of it. Uh but yeah when all of them are searching for gold one become richer by selling a shovel. Yeah I mean they're
not selling shovel right now. Uh it's it's yeah I mean it's a modern day gold for them rights are everything.
They they are just capitalizing on the demand at this point in time. It'll get commoditized very soon. It'll get
commoditized. I think all of this will die down in the next uh especially the the cost factor right uh this will die
down u but what will remain is adoption so what I mean if anything the cloud provider should simply be playing the
role of uh for for adoption the hardware providers if they are trying to play the game of money here
they'll unfortunately lose in the longer run because if not Nvidia Intel will make money intel Intel's playing catchup
here in this case because Intel doesn't really have a lot of u uh footing in in the AI space. Um they do have some of
their systems. Uh but uh it's very very early for I mean I wouldn't say early but they don't have a huge partnership
in this particular space. Um is AI more like a hype? Absolutely not. That's is one thing I want to I want to clear this
is not a hype by any means because AI unlike unlike blockchain is not something that has come in the last
two years, three years, five years. AI has been there since 70 years now. It is it is a massive leap in this space which
is why you start to see a lot more u progress happening suddenly out of nowhere. Otherwise uh otherwise you
would not necessarily observe this kind of a this kind of uh sudden discussions in this particular
space. It's been there for 70 years now. This is not by no means a hype. The current setup is a bit of a uh
I would say the current situation that the market is in is in a bit of a you know everybody's running uh like
headless chickens to try and see what they can do in this. But the reality is if you want to stay invested for you
need to kind of stay invested for long in this right don't don't don't make short-term investments short-term gains.
uh don't try to optimize for that. That's that's the only thing. Uh I I read a very very interesting quote the
other day. Uh there's a concept called u Amara's law. Uh Google it. Uh what what Amara's law says is you know the
industry has this tendency to overestimate the impact of a piece of technology in the near in the short run.
Right? industry has the tendency to overestimate the impact of a particular piece of technology in the short run and
underestimate the impact of it in the longer run. Um so if you want to really make make something out of this, if you
really want your company to stay ahead, stay invested for long. Don't try to optimize for short-term gains. Um so the
best way to do that is build good foundations, right? So build build the right kind of skill set. So if you have
an organization, if you're running a company or if you are let's say in a leadership role and if you are trying to
let's say um guide people on making the right kind of decisions, you need to sort of stay invested in that. I know
not all of us who are here are probably in that space, but you need to think of how do you gain in this in this game.
You need to you need to just stay invested um make the right kind of investments um by building the right
skill sets. That's that's the most important thing. know how the space is evolving. Um some of some of these
things will change over the years. Um technology might change but the core concept will not change. Um transformer
models have come in. Um so transformers will remain the same for the next few years but how transformers are going to
get used to build models and how those models will get hosted that will all change. Hardware will become cheaper,
transformer architectures will still remain the same but but hardware will become cheaper. cloud providers will
make it easier for people to access these models. Fine-tuning these models will become much faster. All of that
developer community, all of that uh you know ecosystem around this piece of tech will change very very rapidly. So long
story short, my point is this is the current value chain. Uh here's how different players are sort of getting
themselves uh in in invested in this. I can share more material on uh how some of the others have um made more money
where the data center lies. data center is an underground facility in a lot of countries um where it's like a massive
warehouse underground um temperature controlled like crazy crazy places but yeah in
multiple areas. Now let's talk a little bit about um you know um some of the more technical aspects of this. I know I
think we've spoken about a lot of these um you know business sort of concepts but let's
let's go one level lower. Let's get into the technical detail. So um so we said we we said generative AI can do all of
these things right. So the origi models u you know can do any kind of um you know let's be more specific here right?
So it can do text generation, code generation, image or or image generation and here I'm also going to say question
answering um video generation and so on and so forth. Right? These are some of the very
very popular applications of generative AI and we said generative AI uh or the genai models. The foundation of the
generative AI models is the transformer architectures. So let's start with some very simplistic
application. Right? So I I'll show you how you can use um or how you can do some of these tasks. Right? uh very
simple examples to start with and then what I'm going to do is then I'm going to go towards the technical you know
we'll go slowly into the more conceptual technical detail of the transformer architecture itself. Okay. So um what
I'm going to talk about is when we talk about geni I said these are some of the applications. So let's let's actually
see how you could how you could do it. Let's let's cut to the chase, right? Without actually spending too much time
on the technical detail, let's let's get to the some of these examples. Let me show you some examples here. So, so I'm
going to set you up for a couple of uh pieces here. Um I'm going to introduce you all to um I'm going to be using for
this example, I'm going to be using the OpenAI models, right? So, the OpenAI models is the Genai models is what I'm
going to be using. Specifically, any of the GPT family of models is what I'll be using. Uh as I said there are many many
many many different models that are available out there. Um OpenAI's GPT4 O GPT4 are by and large the most popular
models that are out there today. One of the most effective models that are out there as well. So I'm going to be using
one of these models for the moment. Number one. Number two, um how do you use these models? There are different
ways of accessing these models. What are the ways of accessing these models as a vanilla user right meaning like an like
an end consumer right if you are let's say bearing the hat of a user then you can access these models to through the
chat GPT interface um so the chat GPT interface is essentially like a chat box right where you can ask any question you
want and it'll generate a response back but you are wearing let's say the hat of a developer
right? Where you want to build applications, AI applications, right? So, you're not just trying to just chat
with AI, but you would want to also, you know, use this to maybe automate some some things behind the scenes. You want
to embed some level of AI capabilities into your existing application, existing code or whatever. Then what you will get
access to is the OpenAI APIs. So, you have an API layer. OpenAI has hosted this model. If you go back here, so the
model that has been built by OpenAI has been hosted on Azure. The model is available on Azure for people to use.
Um, so I will just go to the OpenAI website and I will create an instance of this. I will share that with you as
well. Uh, for this, you know, if you want to try out some things, you can try it out. But just be a little judicious.
I'm going to be sharing my key with all of you. Um, so be a little judicious. my request. Let me show it to you right
quickly how you could do it. Go back the OpenAI API. Let's start with the OpenAI interface itself.
Uh OpenAI. So, so you will have to of course go to platform.openai.com. So, this is the OpenAI platform. I'm
just going to quickly log in. So, I have the platform.openai.com openai.com over here. Let me go to playground.
Let's go to the dashboard. Uh in here, if you see, I have something called as API keys,
right? I can quickly create a new key. Uh I might already have a project. Yeah. So, I have an API key here. I can
also look at my usage uh for the keys that I have currently have so far.
Um the total usage is 26 right this is the total amount of usage um that that my this thing has gone through so far um
and let me just go back here I can actually quickly create a key I'm going to delete the older one
going to revoke this uh create a new key let's create a intellipad at ji test.
Um, perfect. So, this is the key that I now have access to. So, I have this um
interface. Let me just ensure I'm not perfect. Okay, cool. So, now I have my key. I got
the key from here. I have stored my key in av file. I've created a file called env. And I've stored the stored this
particular key in this particular file over here called ent. Now what am I going to do? So how do I access the
model? So the models are available here in the openi website. The models are accessible. I have the key to access the
model as well. But how do I access these models? The beautiful part of this is if you go here, if you go into API
reference and if you see how you can access these models. The interesting thing to install
the official Python binding run the following command. So you need to install this library called pip install
open AI. So open AAI has also created a Python library for all of us to access. They also have a node library as well.
If you want if you're a NodeJS, if you use NodeJS, you could also install it through NodeJS. Um but uh we use Python.
So all that you need to do is you just have to install pip the openi u library. So let's actually go back hereation
pip install openai and that should immediately install the openi library for me. So why is the openi library
important? The OpenAI library is important because the OpenAI library will give me some functions that I can
interact with the model that is hosted on uh OpenAI's platform. But who will give me the the keys, right? Can anyone
access the model? Well, anybody can access the model, but you need the key to access that model. And the key is
what I have given you here. See, remember OpenAI models are not free to use. OpenAI models are closed source
which means they will charge you for accessing these models. It's not it's not for free. It is expensive meaning
rather they charge you but it's not very expensive. It is cheap. It I mean I wouldn't say I wouldn't use the word
cheap but they are reasonable in terms of their cost. For the purposes of this particular session I can give you the
you know I can give you the code and I can give you the the key. You can try it out. Um
so the cost is here. You can actually look at the costing of it somewhere available. That should be available
here. Let me actually show it to you. Uh models. Oh, sorry. This is the API interface for your docs.
Models. So yeah, here is all there is. Yeah. So here is the pricing of these models. So for example, if you take any
of the models and then you say uh this is the images of course uh if you take any of these models GPT40
um $5 per 1 million tokens right um so if you have 1 million words or tokens the cost is 5 million. If you
take for example an older model, if you take uh the GPT40 mini model, the cost is half the half what it is like it's
it's 0.15. The GPT40 is a slightly smaller model, most costefficient small model that is smarter and cheaper than
GPT3.5. So 40 mini is actually pretty cheap, fairly cheap for us to use uh for sure.
So when I actually show you the examples, I'm actually going to you know this is like 0.15 per 1 million tokens.
What is a token? You all have spoken about must have discussed about tokenization
when you might have discussed uh uh text. Yeah, tokenization is what? Tokenization is a tokenization is the
process of breaking a piece of text down into individual words or not necessarily root words but their parts of the words.
A good way to think of it is you know what a good approximation is consider around 75 words as 100 tokens right 75
actual English words would be 100 tokens that's uh very much uh that's how you would pro pro you know possibly use it I
mean like so that's that's like a good so if you were to kind of convert this into a thousand tokens what you're
saying is approximately around 700 words. 700 words is like a word document, right? A simple word document.
If you were to get create a word document using GPT40 mini, it would cost you so much. That's fairly cheap.
0.000015 tokens, which is very very cheap. So, we know that you can use the GPT40
mini model. But the question is, okay, how do you how do you actually access it? So, let's go to quick start. So,
here you go. You want to generate some text. What do you do? import open AI. Uh this is Node, I would assume it's for
Python. Yeah, that's it. So, import open AI. You create a client. Um and then you just simply use this to fire a question.
client.comp completions.create whatever model you want and then you can simply fire a question. You will have to
just tell it what you want it to. So, let me show you exactly how you could do it. Let me copy the four mini model. Let
me go back here. Perfect. So, so here you go. So, I am just going to load the environment. So, when I call
the dotloadad environment, what will happen is this env file that you have this variable openi key will be loaded
into my memory will be loaded into my memory. Now, all that I need to do is I just have to fire
this question. So, client chat dotcomp completions I'm creating an object here. This is VS
Code, Visual Studio Code. Um, you can do it in Jupyter notebook as well. This is a Jupyter notebook in VS
in Visual Studio. You can do it in collab, you can do it in Jupyter, you can do it in whatever you want. It's a
Python interface. Doesn't matter. Wherever you want, you can do it.
Um, so model equal to GPT4 mini. Okay. and I'm saying ro. So remember whenever you're accessing some of these
models right you have to give it you have to tell it sorry you have to tell it saying hey look
um I'm giving it two roles here um there's always two people or rather three people or three entities over here
that are at play one is system so when I say system a system instruction ction is like I'm giving it a persona. I'm
saying, "Hey, here's who you are. Whatever you do, you will have to do it with this persona." So, what is a
persona that I'm saying? You're a writer at a tech blog. Keep the responses short and engaging. Include very contemporary
examples for the questions asked. Right? So, I'm saying you're a writer at a tech blog. Keep the responses short
and engaging. include quirky comments
in the response. Okay. And I'm saying in and include very contemporary examples of the question of
you know examples of for the Okay, perfect. Cool. Um, the user is I am let's say asking it
a question. So the user is me. I am the user or rather this client is the user. So for example, I am asking it a
question. I'm saying hey, what are the differences between EI and JI? That's the question that I'm asking. And this
particular client right now is bearing this persona. The persona is you are a writer at a tech blog. You have to write
it with a certain fashion. So whenever I ask this question, it'll take it'll bear this particular persona while it
responding back. So take a look at this. What went wrong? Uh incorrect API key provided.
Okay. Perfect. There you go. So, um that's the response that it came up with when I
asked it to do it. So, what does it do? Uh let's let me just go step by step again. I made a small mistake. So,
apologies for that. So, of course, I have my user key here. Um and then I've loaded the environment
uh as you can see. And then now here I'm trying to of course access the underlying GPT4 mini model. Um and then
I'm saying u ro user system always remember there are two to three roles. The third role is that of an AI but for
the moment let's forget that um the system role is to provide some kind of a persona for the openi model. So the
model is bearing this persona. It is saying look and I'm telling it you are a writer at a tech blog keep the responses
short and engaging include quirky comments in the response and blah blah blah. Uh and then what I'm also saying
is now I'm asking the question as a user. So let me actually reorganize this a little so that it looks uh logical
because it might seem like I'm asking the question and then so here you go. So that's the role the system role and then
the role the as a user I'm asking this particular question and then I'm asking it to generate the response and the
response as you can see is here absolutely let's dive into the exciting world of AI and JI here is uh what it's
saying AI artificial intelligence uh is the broad umbrella under which all kinds of smart technology fall think of it as
a wizard that can do many tricks everything from recognizing ing your face on Instagram to analyzing stock
market trends. Basically, it's like a super intelligent friend who can ace trivia night but might struggle with
creative writing, right? No shade. Um, Genai on the other hand is is cool cousin of the AI family. So, why is it
kind of coming up with stuff like this saying is a cool cousin of the AI family who's not just smart but artistic too.
is designed to create new content like generating images, music, blah blah blah. Picture chat, GPT, and Dolly as
your artsy friends who are at a party who doodle the wildest designs and web poetics on its while sharing memes. So
is actually able to write something like this specifically because I've asked it to include quirky comments and I've
asked it to include contemporary examples. Uh, exactly. This response is completely
generated by Genai. Now if you look at the examples if an AI example if an AI can analyze recommend your next Netflix
binge thanks algorithms. Genai can actually imagine the whole new movie script and invent an entirely new
character to spice things up. Why settle for another romcom where Jai can throw in time traveling uh cat as a
protagonist? So so you see the and it actually added these memes as well. It added a cat. It added a a rock, you
know, sort of a rocket meme over here kind of just to say a time traveling cat protagonist. Um, in short, while AI is
your reliable assistant, Jana is an is an imaginative storyteller. Uh, both have their perks, but definitely one has
more flare. Um, keep an eye for both of these techno wizards who know that they'll conjure up what what they'll
conjure up next. So, again, a super cool way of explaining what AI geni is. all of it just because I've I've given it
this kind of a persona. Now, let's change this up a little. Let's let's say you're a writer at um now I'm going to
say you're a writer at um economic times responses uh you're a
sponsor formal um and
professional include contemporary I'll just I'll just put that right. Um,
now I ask it to do this. Let's see what it comes up with. I would expect something super boring.
There you go. Artificial intelligence refer to the broader field of computer science
focused on creative creating systems that can perform tasks typically requiring human intelligence such as
reasoning, learning and problem solving. Genai on the other hand is a subset of AI specifically designed to generate new
content or data such as text blah blah blah based on your input and whatever and then they say in summary all geni is
AI but not all AI is geni. Yeah, I mean does the job. The thing is this doesn't have a persona as it as it said. It's
got a lot of flare because it actually is created by AI with a specific uh you know it has its own persona. It's kind
of spicing it up a little. Uh it's talking to let's say a teenage uh um you know a teenage individual as you can
imagine it kind of resonates with that particular person. So you can kind of give it these personas. Um and that is
what you mean by system um you know messaging over here. Now let's go one step further. What else can
it do? Right? If if it has generated content, what else do you do using chat GPT? What else do you do using chart
GPT? You of course ask it to write code. Of course. Sure. Well, let's let's look at code as well. Let's come to code in a
minute. Uh I actually asked it to write a poem here. I'm saying write me a small poem or rhyme about geni and then the
system prompt is you're a school teacher for a fifth grade student. Uh and then I asked it to write a poem. Let's go.
Here's here's the poem. In the world of tech so bright and gr grand, generative AI leads a helping hand. It crafts new
stories, draws with flare, creates new worlds from pixels and air. Uh it learns from data young and old. A mind of
suckets truly bold. Like kickass. I mean this is this is very good. I it's also it's also kind of leaving in a little
bit of thing here right? So dream with tech but don't forget the world needs your passion death. Uh generative AI is
here to stay but it's you who who it's you who leads the way. I think this is this is this is amazing. It's also kind
of leaving that thing in right it's also telling you not to fear. Now what I could also get it to do and and this is
the part that I was talking about uh with the others is
um I can also get it to write code for me. Um of course you could do the same using chat GPT as an interface as well.
But again here is where I want you to think of an interface within your organization where you just build and
this is by the way one of the products that I'm building with my company right where for developers for business
analysts my business analysts they spend a lot of time trying to analyze and trying to analyze structured data. So
how do I then provide an interface for all of my analysts in my company where they can simply say give me
you know summarize the sales in so and so particular market or summarize uh the sentiment of so and so particular market
and so on and so forth um then that's what I'm asking it to do here write a SQL query so how if I say summarize the
sales remember that genai is a language model or these are large language models they understand language. They will not
understand numbers just by themselves. So what you need to do is you need to somehow
get them to understand language or get them to understand these numbers. But one way to do it is you say okay keep
the data where it is. What open or geni models are very good at is
generating code. So I say you know what you generate the code and then you use that code to ex to to query against the
database extract the response and summarize it. So that's what I'm asking you to do. So one atomic actish action
in that whole exercise is to write a SQL query. I'm saying you're a data analyst in a technology company. You're at high
quality bugfree code and your expertise is in Python and SQL. Um and then I say ensure that you only return a SQL query
or a Python code and nothing else. The response can be as a string or a JSON. So this is one of the many activities
over here, right? Imagine there are multiple other such activities that this I can create
multiple such agents that could do multiple such actions. So when I ask a question, hey summarize the sales for
me. What I'm asking it to basically do, one of the actions that I'm asking to do is saying write a SQL query to analyze
sales of each of the stores in Europe. You have access to sales database and customer demographics. So when I execute
this, look what it does. It goes ahead and actually creates the SQL query for me over here. Then I can use the SQL
query to go ahead and query it against the actual database. Get the response. Then pass that response back again to
Chad GPT. And I say you know what go ahead and summarize this information for me
and then it summarizes it. How does it know what table it contains? So then what I can do is I can provide that
information also over here. Currently I just mentioned you have access to sales database and customer demographics. But
what you could possibly do is you could also pass the table descriptions, column descriptions, all of that into this and
you can generate a response. So you could actually pass that as additional context over here and you can get it get
it to generate the response. Yeah, it can be different in each system, but there is always a way for
you to extract it, right? So you can always in a in a given DB, you will know all the tables, you'll know all the
columns, the table description, column descriptions would all be available for you. So you should be able to simply
query again. So if it's not there, you'll have to fix that. But here's another another piece of thing that you
could also do with with uh with with the OpenAI, which is this OpenAI model. I'm saying clients.generate.
Now I'm actually generating an image here. A coder underwater sipping a coffee and I'm asking it to generate
this image for me uh and I'm saying I want a 1024 x1024 HD quality image. Um and it's going to return a URL for me.
Um and once I click on that particular URL I should be able to access that particular image as well. Let me the
URL. There you go. If you see blob.core.windows.net net which is essentially like a you know
Azure capabilities. There you go. Let's return the coder underwater sipping a coffee. There's a coder underwater
a coffee. What I could do is I can do this um this as well. Create a fun ad for my cola
beverage brand. It's party uh it's a party environment in the background. Um I can say it's a party uh you know
environment in the background focus on condensed droplets on the can and I can say I don't know the can is blue in
color or rather let's go is um is teal in color and uh
and a portion of the can is also transparent. parent with blue with uh
pink liquid inside. I don't know, man. I'm just coming up with something. Let's see. I'm being
creative here. Let's see if this is going to be equally creative or not. Genai is the concept. OpenAI is the is
the company that's behind it, which is true. And uh GPT is the model that is enabling
it. Uh almost there. But yeah, kind of it kind of added the teal and the pink, but it didn't kind of make it a
transparent bottle, but it did the other stuff. As you can see, it did the other stuff. Um but I
can ask it to create uh a realistic image, right? So I can ask it to create the
image as realistic as possible. So it will hopefully create a a realistic image not not that kind of a
looks like a very cartoony kind of image. Let's see if it creates any different
kind of again I'm not very very happy with this but it does have the droplets that it is
focusing on. Um, I could actually provide um
a I can try to give it copy im copy
image link. Let's see if it takes us to that image. Very good. Let's see if it creates something like
this. Let's ask it to create it something like this.
Uh I'm just going to remove all of this. Can we ask for regenerate if you don't? Yeah, just reexecute it. Regenerate.
That's it. Can it generate 3D images? Yes, you can get it to generate 3D images.
Use glass bottle in the prompt. Yeah, I mean let's read it.
Yeah, it it did bring the Coca-Cola thing, but of course it doesn't use the actual branding itself as you can
imagine. Uh but it did bring the Coca-Cola kind of bottle design as you can see here. It has taken some
inspiration from this. But the thing is it will not use of course the absolute branding of Coca-Cola straight away. Uh
you'll have to force it to do it. But cool. Awesome. Um so guys uh I hope you get the idea of of this right. So
this is let me just see accessing the Sora model. I'm not sure if the Sora models are available for public
consumption. Uh but let's go here. Where are the models?
uh API reference there's audio yeah the video modules are not available
to what I know you you need to access it through the interface I'm not wrong but let me check by the way you you could do
this as well right create image variation so you can pass an original image and you can ask it to create
variations of it you can pass any of the existing image and you can create variations of it as well. It's kind of
kind of kick because you could essentially ask it to create multiple types of the same image uh in some
sense. Uh it it's fun though. Uh right. Um and then yeah the video the Sora models
are not available here. access Sora through API.
There is currently no way to access Sora from a website or an API. So there's no way to do that. So as a Sora is not
available. The video is not available. You can access it through the through the front end if you want to through the
you can get chat GPT plus and then you can you can do it from the front end if you wish to do that. Okay, cool. Um
let's actually try this last one which is the variation. I'm keen on understanding how that works.
Um open AI and uh
let's try one of the existing. By the way, just to let you know my uh just in case you are interested.
So this my friends um Beex is one of the brands that I that that my company uh has
built. Okay. Um so what we were able to do is we kind of launched this product called Beex
Autonomous. Um and this was built on top using midjourney.
So this product called beex autonomous by the way it's in the market right now. Uh it has actually launched in the you
know people is actually now available for people to actually consume. The product itself is available uh for
people to kind of consume or whatever. Point is from the recipe to ads to bottling to everything of this product
has been made using geni completely. That's the product. It's flexon. Uh everything has been made by geni.
It's kind of it by the way it's sold out right now. It's not available but uh they've only launched like a few
versions of it. Uh this is something that uh was uh that was created by the company that I work for.
Um let's go for the images. Let me find one of the
link. Let's see how this work. It needs access from the local machine.
Yeah. So I can ask it to actually create uh variations of this.
go back here and um what happened?
Uh it has to be a PNG. Okay. It's mandatory for it to be a PNG. Okay. Okay, let's see
part. Let's see what it does. I can of course provide it more prom.
Yeah. So I see some question some points here. Advertising agencies should be clearing their lives now
agency modeling. Yeah. So uh so I'll tell you how advertising companies are yeah it
yeah kind of boring though but it kind of created a version of it. uh yeah not not very not very pleasing
so to say but I can of course write a prompt u and I can ask it to specifically
operate a certain way of course over here um I can I can guide it in a specific direction of course I can add a
certain prompt and I can get it to do a few things stuff like that but anyways uh point is um I I see a point there
about how agencies are using Guys, this look um this is where A&Z are of course not going to be using
this through through coding interfaces. But what's happening with marketing agencies is marketing agencies now they
of course use tools like uh your um Adobe Photoshop uh or let's say Figma and stuff like that and what's happening
right now is these these capabilities are now coming as a part of that right so these integrations are coming as a
part of Adobe so Adobe right in Adob Photoshop if you take a premium version Adob Photoshop has launched something
called as Adob Firefly um And Adob Firefly is exactly the same thing as what you're currently seeing on the
screen. So Adob Firefly is essentially the same thing. It's a it's generative AI for creatives.
So you could kind of do exactly what we just did. So you could do generator fill image generator
uh blah blah blah. You can do like a bunch of different things uh online uh with uh with stuff like this. Super
cool. Very fast. So as you can see like seems like a Pokemon only this one. I don't know what that is. So yeah. So
point is they're already using this extensively in their their work. Um code
I'm not sure if you all have heard of GitHub copilot. It is already a part of you know these tools have already come
in. GitHub copilot is already a capability that kind of provides you an interface where you could automatically
start writing code. um with writing. I don't know if you've heard of Microsoft Copilot. Um Microsoft Copilot is a
capability that gets integrated straight into your uh word documents, your PowerPoints. So you can actually ask it
to write content for you, write a document for you, um write an email for you for that matter, summarize emails
for you. Important something that I do use quite extensively. I actually use Microsoft Copilot very very extensively
for um rephrasing emails. Right. So yesterday I was asked to write a business case to explain why I should
continue hiring in my team. I just wrote two lines on chat on on sorry on Microsoft copilot and I said like hey go
ahead and write write this email for me. It actually ended up writing it in two minutes. Um and I'm done. Otherwise I
would have had to spend like half an hour writing that complete business case. It would have been an absolute
waste of my time. Um so stuff like this super super easy uh to do with uh capabilities like uh like Jenny. I don't
write one email without rephrasing using Microsoft copilot. I don't write one memo or a document without using
Microsoft copilot or or for that matter even chat dbt sort of thing. So these are how it can impact you on a regular
basis with your work. But then what you can do my friends is you can take this power of these capabilities not just use
it for personal productivity but actually take it one notch up. Right? You can combine these capabilities
multiple ways and then create agents that can automate workflows that you can combine these capabilities together to
let's say from internet extract from SQL database extract read from a PDF document combine all of it summarize and
write a report out or send out email all of this in a single shot which would have been super complex earlier. you
would not have been able to do something like this and all of this without you having to tell it what to do. You can
just provide it those capabilities. You can simply write a question and it will automatically do that for you one after
the other be super super um you know capable when when when things like these start happening. I'm going to touch upon
things specifically around how the um transformers models work or some of these generative AI models actually work
right. Um so if you remember we spoke about um how generative AI models are based on these core concept called as
the transformer architecture. So I'm going to touch upon the transformer architecture a little. you know this was
not discussed with all of you um so I'm happy to redo it um so we can we can touch upon transformer architectures a
little and that would already should set us up well now it's going to be a bit of an intense piece of next few minutes um
or or next maybe half and half 45 minutes everyone wherein we're going to delve into fair amount of detail with
regards to the transformer architectures one thing I want you all to know is that this is a bit of a complex setup. It is
a complex architecture. Um but we will discuss it anyways. Nevertheless, um and then from there once we at least
understand you know thousand ft um high if you're able to understand how this works then we can get into uh the actual
detail itself then we can at least move on and then we can talk about some of the other concepts uh about this. Okay.
Um so for the first part of our today's session, we're going to focus as much as possible on the architecture of uh the
highle architecture of um these models of the transformer architecture. Um with that context, let me go straight in.
Sorry again. If there's one thing that you might have must have that we all should
have learned by now is that look this this space is evolving fundamentally if you if I I'm assuming
you all have used chat GPT if you can do things like charge GP if you can do what charge GPT is doing through the
interface if you know what it could do potentially you all can build those kind of capabilities behind the scenes as
well right so using the APIs you could also do all of that but if you want to build a PPT If you want to build
marketing content, certain things are easy, certain things are not very certain are certain things are slightly
more simpler to do. Certain things will require a lot more software engineering because you might have to interact with
PowerPoint, you would need connectivity with Outlook, you might need connectivity with 0365 suite, you might
need connectivity with uh a bunch of other things um and stuff like that. So my point is
all of that is possible. Um I think what we will be focused on to start with is to understand and I'll tell you
something as well right so the the the pieces of example that you're talking about
everyone the pieces of uh examples that you are talking about like using it for powerpoints using see these
things will get automated right somebody or the other will come and they will try and make it you know like Microsoft will
just do today they might charge it, tomorrow they might make it free.
Um, so point is that it'll become super obvious. I I'll give you one example, right? Biometrics,
right? So let's say face recognition, face ID on your mobile phones or let's say your fingerprint scanners, that's
all AI. That's all computer vision. But nobody calls them as AI today because it's just there. it's so
commoditized that everybody has access to it and these companies are just using that piece of technology and they're
just embedding that into the products. So the PowerPoint thing is exactly going to become that. um what we should sort
of be looking at is not the whole you know how can I use it for powerpoints how can I use it for word documents
instead of looking at that you should look at okay how can I use this piece for let's say
maybe automating workflows how can I automate uh how can I use generative AI or generative AI models to let's say
respond back to query customer service uh quest you know your question how can I use this generative AI models to let's
say automatically how can I build agents that can automat automate workflows. So I think you
should sort of look slightly more broader not just uh with with those smaller quick wins but again we we'll
get there. I think we'll slowly once you start doing a couple of examples we'll also get there um super easy to do all
of that. I'll just show you the what I'm going to show you is I'm going to give talk about a couple of tooling and then
some of it will become super easy for you to do. um for for some there are tools that are
already available as just letting you know. So let's let's get into a little bit of detail on the transformer
architecture itself right um as I said the most fundamental part of let's say any of these GPT models um if you talk
take for example is the word GPT right so
when we talk about the word GPT GPT stands for GPT is just one of the many models
generative pre-trained transformer that is what GPT stands for
that is where the GP and the T sort of come from right generative pre-trained transformer
that is what we mean by GT right now these GPT models are the sort of models are again there are many kinds
of models but I'm taking GPT as one example. So when you talk about GPT3, GPT4, Chad GPT, they're all of the
family of they all from the family of transformer models, right? So what are these transformer
models? What are transformer architectures? Right? Uh
now to be very specific right so the transformer architectures um just a second.
So the transformer models have sort of been introduced by
this paper called as attention is all you need. Now this was the paper that was first published in 200 I would say
um approximately in um 2017 is when it was first published. Um and since 2017 this paper sort of went through a couple
of re you know couple of revisions of course but this paper is the paper that kind of
made that that that sort of uh made so that kind of transformed was like a game changer this particular paper. All
right. Um what is it about this particular paper? So if you look at uh the abstract on this paper right I'm not
going to go into too much detail I'll start with here right so the dominant sequence
transduction models are based on complex recurrent or convolutional neural networks that include
an encoder and a decoder okay now I'm not sure uh where you all did you all discuss sequence to sequence
models or uh if you have learned ls TMG would have also discussed the N the encoder decoder models. If not that's
okay. I'll just briefly touch upon it. Um the point is that the the state-of-the-art
let's say translation models you you take translation translator
um or any of these um state-of-the-art models. Now this was in 2017 that I'm talking about. They involve either a
very complex recurrent neural network or an LSTM so to say. The best performing models also connect the encoder and
decoder through an attention mechanism. So there was a mechanism called as an attention mechanism as attention
mechanism which was introduced before 2017. Now there is a new simple network
architecture called the transformer based solely on attention mechanisms dispensing with recurrence and call
basically getting rid of the whole idea of recurrence and convolutions entirely. So again if you remember yesterday I
spoke about the fact that when you learn about recurrent neural networks the reason recurrent neural networks are
firstly the way recurrent neural networks work is you recursively pass let's say one word after the other and
then you try to predict the next word. In this process um you are sort of trying to learn
um the probability of the next word bases the current word and then there is one weight matrix that you have which
you're trying to recursively learn over time. Right? So you're you're saying your sentence is a sequence of numbers
or a sequence of words and one way to capture the dependency of one word to another word is by going from right to
left and thereby capturing uh you know trying to use the first word to try and predict the next one and so on and so
forth. So that's how the whole recurrent neuron network sort of works but as I said it's super slow. Um so experiments
and what they are saying is we've gotten rid of these network architectures that have to do with recurrence or even
convolutions for that matter. We've introduced a new model called as transformer architecture and this
transformer architecture is solely based on the concept of attention. So we have to somewhere learn what
attention is to start with. We have to understand what is attention. So we'll discuss about attention in a few
minutes. Um and this attention mechanism um gives us a very very good understanding of uh firstly we'll learn
about the attention mechanism then we will learn about the transformer architecture itself. Okay. So what this
paper is saying is look so far all of the state-of-the-art models have been built using trans uh have been built
using any of the um have been built sort of using any of the
um u recurren recurren rec recurrent neural networks or convolutional neural networks. We're introducing the concept
of attention. Along with the concept of attention we're also proposing this idea of transformer architectures. These
transformer architectures have nothing to do with u they're basically chucking the whole idea of of uh um of recursive
recurrent neural networks or convolutional neural networks. They said these are great but these are useless.
Let's simply get rid of them. We will talk about something completely new and these are now going to collectively help
us build this architecture called as transformer architecture. Okay. Um and this transform architecture our model
achieves 28.4 blue with a WMD. Uh yeah so these are some scores. Blue score uh is a is very very good score for uh um
for content creation content generation. Um and they say that this particular model has sort of outperformed some of
the other models. They will also talk about if you look at this and I I'll specifically go into the more I'm not
going to walk you through the paper but I just want to touch upon some of these concepts here. Recurrent neural networks
LSTMs and gated are you know gated recurrent neural networks in particular have been firmly established as the
state-of-the-art approaches in sequence modeling and transduction problems such as language modeling and machine
translation. What do you mean by language models? Language models are models that are always trying to predict
the next word. You have a sentence, you have a sentence, you pass the first word into
the model and try to predict the next word. Those models are basically referred to as language model. Um, so
again, the sentence says it for itself. NLP. Exactly. This is all natural language processing only through all of
this is NLP. We're talking about text. We're talking about natural language processing. We're talking about text
itself. Right? Um numerous e efforts have since continued to push the boundaries of
recurrent language models and encoder decoder architectures. Again lot of stuff has gone into this in 2017 almost
until very recently also right lot of the machine translation when you talk about machine translation we're talking
about language translation um models that you see they've all been predominantly
um based on this the recurrent neural networks transform and and also more specifically these encoder decoder
models sequence to sequence models But unfortunately recurrent neural recurrent models typically factor computation
along the symbol positions of the input and output sequences. So there is basically
aligning the positions to steps in computation time. They generate a sequence of hidden states. Yeah, again
we don't need to get into too much detail but the point is this inherently sequential nature precludes
paralization. So it kind of lets us it doesn't help us with parallelization which becomes critical at longer
sequence length. So when you have large sentences it becomes very very complex as memory constraints limit batching
across examples. So like you cannot have large sentences to deal with. LSTMs try to handle for it but again
computationally they become very very fast. Recent work has achieved significant improved in computational
efficiencies through factorization tricks and conditional computation. There like some specific tricks that
have been put in place but still the fundamental problem still remains right. It's improved but it's still not the
best solution. Attention mechanisms have become an integral part of compelling sequence modeling and construction
models and various tasks allowing modeling for of dependencies without regard of their distance in the input or
output sequence. Again, long story short, attention mechanisms were brought in. Attention mechanism was very good.
We don't know what attention mechanism is. We learn about it. But again, this is the premise. Attention mechanism by
this paper was already introduced. But they were saying that this attention mechanisms were always used in
collaboration with recurrent networks the RNN. So we have to understand what attention
mechanism is somewhere we need to talk about it. We'll talk about it in a few minutes. Um but then they're saying even
though you bring in attention that's not that did not solve the problem because they're still working with RNNs and
RNN's fundamentally have the problem of time. You cannot paralyze it beyond a certain point. In this work, we propose
the transformer, a model architecture suing recurrence, basically getting rid of uh recurrence and instead relying
entirely on an attention mechanism to draw global dependencies between input and output. So somehow they've gotten
rid of the idea of learning the language sequentially um and just treating it as some way to
extract all of this information together, right? you're not learning it anymore sequentially. There's like one
very interesting way now to capture all of this or you know in parallel um and and getting rid of the whole idea of
sequence transformer allows for significantly more paralization uh and can reach new state-of-the-art in
translation quality after being trained for as little as 12 hours on eight P100 GPUs. So just training it for eight
hours on uh you know in just eight GPUs was able to was able to outperform some of the
other models that were state-of-the-art at that point in time and now your GP2 models GPT models my friends are are
really really big right so this far bigger than what this so that's the background okay so that's the background
here I mean I'm not going to go into the paper itself but do you understand some of the challenges here like broadly of
course there's some some things that we probably don't fully understand but that's okay but broadly do you
understand the setup here like why this architecture was brought in to start with and what is the biggest advantage
of a transformer architecture as opposed to let's say um any of the existing state-of-the-art architectures now let's
delve into a little more detail so there are things that have been spoken about over here in the just these three four
paragraphs now my friend here is a thing and and I should not sound very preachy here but without sounding too preachy my
ask with all of you is see if you can spend some time in reading such papers try to I would not say I am reading it
myself probably I I I try to read it but see the moment you read such papers it it opens up a complete Pandora's box
because now you're like okay there's so many things that have been talking about in couple of lines that you probably
have not even heard And remember you've been learning and training on AI for the last
for the last two three months or maybe even more than that some of you. So thing is there are so many topics that
have been discussed in these three four paragraphs that you probably haven't even heard of so far. So again or most
of them you have there are some that you haven't even heard of. A good um a good way to keep your understanding
in check. Uh just letting you know that uh these are some of the areas where you might you could possibly get lost a
little. Anyways, um it's it's important to read papers. That's that's all that I'm trying to let
you know. Um any website uh or Yeah. So there are a lot of websites my
friend. Um one of the most popular website is this paper called is this website called papers with code.
Uh it's sort of the latest and greatest as far as uh you know machine learning is concerned. You typically see on this
uh you have the papers, you have the the code. There's like long form articles as well. You can read about it and stuff
like that. So lot a lot of lot of interesting piece research that happens in this particular it's just an
aggregation. Yes, it's an aggregation of a lot of these areas. All right, let's go back. So now let's talk about the
concepts that have been just briefly touched upon here. So let's talk about firstly this idea of firstly let's look
at how the transformer architecture looks like. Yeah. So this is my friends the transformer model architecture.
If you see there are two parts to the transformer architecture right on one side everything that you
see see on the left is referred to as an encoder and everything that you see on the right
is referred to as a decoder. So there are some as I said there are a couple of things that we have that has been spoken
about here that we haven't that you may not have heard of heard of. So we'll try to go one by one here
right we'll try to understand each of this one by one there are
things like in when we looked at that paper right there were things for example like
so when we read that paper there were these three relevant topics that were discussed, right? So they said when you
talk about um sequencetosequence models, encoder decoder architectures have sort of become the most popular ones and
within this um the encoder decoder models there was something very interesting that they have been
primarily based on RNN's or LSTMs or even GRUs for that matter. They're typically based on these models. Let's
understand firstly what is an encoder decoder model. I mean I'm going to briefly touch upon the encoder decoder
model and and how it fundamentally works and then we can get into the detail a little. So what is an encoder decoder
architecture? Historically when you talk about tasks like for example let's take the tasks like let's say uh
uh machine translation task. What do you mean by machine translation? Let's say you're trying to translate
from English to Spanish or English to Hindi whatever
that that language is right. So if you take a sentence like this from English to Hindi
um the machine translation is essentially a neural network or machine learning
algorithm that tries to convert um any input that is in English to Hindi. So how do you train a model like
this? So the the the most popular architecture that was used and that is still used
that was previously used as well that is still used is this architecture called an encoder decoder architecture.
How does the encoder decoder architecture work? So the encoder decoder architecture has two parts to
it. the first part. Sorry. So here you go. Let me explain how an encoder decoder architecture works or or encoder
decoder model fundamentally works. So this is a very good example of an encoder decoder model.
Uh let me just try to find a simple example. So let's take for example uh so let's say the sentence that you
would want to translate is um you know the the the sentence here is um how
are you doing? Let's say there's a sentence like this. And now in your encoder what
you're typically going to have is you're going to have the first part all that you're simply
going to have is you're going to have four RNN blocks right this is RNN's now this can be RNN this can be LSTM this
can be GRU doesn't matter four RNN blocks the input into each of this is essentially one of the words so the
input can be how are you
doing? Now remember when I say this is the word, what I'm essentially meaning is not the word itself. It is
essentially the embedding of that particular word. Right? It is the embedding of this particular word that
is going to go as an input here. You will pass this as an input here. And once you pass this as an input, what's
going to happen is of course, right? Once you pass this as an input, what's going to happen is the these in these
out or rather these inputs will then combine themselves, right? These inputs will then sort of
combine themselves in some form or shape. By the way, there is a recurrence here. This is the same RNN. So there
there is you know at time t =0 t = 1 t = 2 t = 3 essentially you're essentially going sequentially here there is that
recurrence over here or see you're going sequentially from left to right and then what you're trying to produce
is as an outcome from here you're trying to produce a vector. you're essentially trying to produce a vector over here
called a state vector or an embedding vector or an encoded vector whatever the vector is. So these inputs that you have
here that is essentially converted into a vector V. Now this vector you call it a state vector or you call it a embedded
vector whatever essentially it's an embedding of the original input sentence that has sort of been created as an
outcome from the RNN. Now this vector is now passed as an input into another RNN. So this part is the encoder. So the
encoder over here has taken the original sentence and it has converted that into a input vector right it has combined the
complete context over here and it has created one vector as an output that vector is now passed as an input into
another RNN. So this is RNN one right and this is passed into another RNN where this RNN
this RNN's only job right this RNN's only job is to take this vector as one of the inputs along with this vector
what it also takes it is it takes another input from here right so like for example
um you sort of simply give a a simple token called start. And what this tries to do is given this particular vector
and given this start token, it now tries to predict the next word or the word in Hindi that should come out. So this
tries to predict the word up and then after this it goes further. Now it goes for the next time step. The same RNN now
takes the word up as an input or it takes actually both the words as the input. It takes this state vector. This
state vector is of course going to be there. But along with this it takes start and it also takes the word up. And
then it tries to pick the next word which is probably chess. Um and then the same RNN again takes
these three words as input, right? It takes all the three words as an input. So it says start and care.
Uh, and then it predicts the next word, which is probably Oh.
And then it continues to do this until it's until it predicts a token called end of sentence. It
continues to do this until it predicts a sentence, a token called end of sentence. The moment it predicts this as
the token, it stops generating. It stops generating. So point is that this piece over here
is essentially taking an encoder right taking a piece of sentence as input. See word by word it is taking one
by one by one word and then converting this into an embedding and then this embedded vector is now being passed as
an input. Just to let you know this embedded vector is passed as an input like this actually.
This is essentially passed across all of the uh you know across all of the decoding uh decoder path.
Now the objective of doing something like this. Yeah. So every sentence every sentence
um you know wherever we look at the sentences right what we typically do with these
sentences is we always start of the sentence and we and end of the sentence we typically add these
additional tokens over here we typically pass the SOS and EOD you know end of the sentence as as
additional tokens on either side just to indicate that this is the start and this is the
So technically speaking even here as well the first word will be embedding of SOS that should be the
first RNN and the last one over here would be embedding of EOS end of sentence. So you sort of pass
all of those as inputs and then you're expecting the other words to be generated as an output and then you're
hopefully you're trying to predict continuously until it predicts SO you know EOS as the output. Embedded vector
is the only input that is considered in the decoder. That is absolutely right. Again the embedded vector is the only
vector the vector V is the only input vector that is being considered in the output
along with of course whatever you have predicted in the previous time step. The only two vectors that are being
considered here are this vector vector V and whatever words that have been predicted. Now this all of that is being
considered as an input. Only these two are being considered as an in. Now why why are these ve why are these now this
remember remember this is this has nothing to do with the transform architecture. This is a very very
popular architecture for doing any kind of translation for doing any sequence to
sequence. Why is this useful? Why do I need to have something like this? You know, I could simply have an RNN like
this. I could simply say I could have a sentence like how are you? Ando,
right? I could simply do a word by word translation. Right? I could simply do a word by word
translation. I could just have one RNN. I can pass how as the input and I can predict another word. Similarly, I can
pass another word into the same RNN, predict another word and so on and so forth. I could do wordby word
translation. Why do I need this kind of a setup? Well, the problem is especially in languages like Hindi,
whatever word you pass as an input, the translation is not always in the same order. number one.
Second thing, the output translation doesn't have to be doesn't have to have the same number of words either. So you
cannot simply make a word by word translation. You cannot just have one RNN predicting one word after the other.
Right? Which is why to be able to overcome that kind of a setup, we say okay, let me take the complete input
sentence, let me embed that complete sentence and then convert that into one vectorzed representation,
one long vector. After that vector has been created, I pass that vector and then I train
another decoder which then word by word tries to predict my outcome, my output. Right? This is the first word. is the
next word is the next word next word and so on and so forth. Right? So that way I'm not forcing my
model to always predict word by word but rather I'm rather I'm not forcing it to exactly predict the translated word from
English to Hindi but rather I'm saying you know what it doesn't matter how many words are there in the input the output
can have as many as words it can and it can be in whichever order it can be that way the sequencetosequence models have
become a lot more preferred choice of translation um than your regular you know RNN
classification simple classification based model. Now that you understand the encoder decoder setup the encoder
broadly what does it do? So I can simply summarize an encoder like this. The encoder decoder architecture can be very
easily summarized this way. encoder decoder. An encoder would have would take the input,
right? And this encoder would generate a a vector V and there is the decoder
which takes this vector as an input and then predicts the output. Right? This is a embedding.
There are multiple names to this. People call it an encoder. Encoded vector. I Okay, let me not call it an embedding.
I'll simply call it as an encoded vector or a context vector what doesn't matter whatever that term is that is an encoder
decoder architecture every now. So you pass an input you have an encoder. This encoder typically is an
RNN model was an RNN model. Then you have a decoder. A decoder is also another RNN model which takes the
spectra as an input and generates the output. Now once you have this set up
what was being spoken about is the fact that you have these models are some kind of an RNN model is very restrictive.
Why? RNN's are time-taking. RNN will only operate sequentially. lot of problems with RNNs. Hence introducing
transformer architectures, introducing something. So two things, one is the problem with the RNN itself,
right? RNN's are slow, cannot be parallelized. The solution to that is the transformer model, the transformer
architecture. The second is this encoded representation is also at
times not very rich. Right? So this encoded representation vector is not rich enough doesn't
capture enough detail of the sentence. Hence the outly the solution to that is attention. There is a technique called
attention which we will discuss right now. Yeah, paralyze means you cannot you cannot pass all the words at the same
time. You'll have to go one word after the other, right? Because that's how language works. You'll have to predict
pass the first word then the next word then the next word then the next word and then create the embedded vector.
Then after that the prediction has to be sequential of course but even the encoding part of it has to happen
sequentially. Pass the first word then the next word then the next word and then the next word. You have to go
sequentially which is why you cannot simply go all observations at the same time or all the words at the same time.
You cannot do everything at the same time. That is what we mean by it is it cannot be parallelized. So the two
challenges one is RNN are slow and cannot be parallelized. Hence the out the the answer to that is transformer
architecture. the encoded vectors are at times not rich enough and hence the outcome or the change that was brought
in is the attention architecture. So this my friends is the answer to this. Now we'll of course have to get into a
lot of detail here but broadly to start with whatever you are seeing here forget about everything else stick with me for
a couple of minutes right whatever you're seeing here forget about everything that you're seeing here this
is input this part that you're seeing here is encoded this is your encoder right so if you were to just simply look
at it what you're saying is you're saying you have a block being passed into an encoder
and this encoder is generating a certain output. This encoder is generating a certain output or a some kind of a
vector. It is generating a vector over here. Now this part my friends don't read too much into what's there inside
it for the moment. This one whatever you're seeing here is simply nothing but the vector that has
been generated is being passed as an is is being passed as an input. This is your decoder. This vector is being
passed as an input and then there is output embeddings which is nothing but the words that you have predicted like
start of sequence and stuff like that. So the outputs and then this is essentially going to
predict the output over. So this is also what do you think this transformer model
is also an encoder decoder architecture. It is very similar to that of your you know this kind of a model. Your
transformer model is also an encoder decoder architecture. You take an input you encode it into some kind of a uh
numeric representation. You take that numeric vector pass it into a decoder and then one by one you generate
outputs. So your decoder is also or rather your transformer is also an encoder decoder
architecture. The only difference is that the stuff inside this it's not an RNN anymore. This is not an RNN anymore.
This is something that is based on top of attention. RNNs are completely thrown thrown out of the window. This is this
detail that you see here. This is not an RNN anymore or a recurrent neural network for that matter. It's based on
something called as attention. Now what is attention? We need to understand what attention is. That is
what we will learn for the next few minutes. We'll understand what attention is. Why do we need a vector here? Okay.
The thing I'll give you a simple example. Okay. I'll give you a very simple example. So imagine you are in so
you traveled let's say to China and now for you you don't understand
Mandarin or Cantonese or any of that any other Chinese dialect you don't understand the languages in China let's
say somebody is speaking to you in Chinese you need to understand it you understand Hindi very well
what do you do you put a translator in between what is this translator doing this guy in between is taking Chinese as
an input in his head. this guy translator taking Chinese as an input converting that into some kind of a
common language or common understanding of this Chinese language converting that into some kind of a
numeric representation numeric numbers are universal right it it has nothing to do with language so you're taking this
Chinese input converting that into some kind of a in his brain or her brain the person is converting converting that
into some kind of a common commonly spoken language or commonly understood language and then you're
taking that commonly understood output or that language and then passing it and then converting that into English.
Right? So all that you need to do is you don't need to then you don't need to always build a Chinese
to English translation. All that you need is okay, can I have an a model that converts this Chinese language into
these numbers and then can I have another model which take these numbers and then converts it into English. So
that way I can put both of these together and I can always do translation very efficiently. So exactly vector this
this vector is simply nothing but numbers is a numeric representation. That's the common grounds that you're
bringing these these two you know distinct languages to because models understand numbers very very well. So I
take Chinese I convert it into numbers through some kind of an encoder and then I take a decoder I take these numbers
and then convert that into English or whatever language I want. So that is the idea of the vector. So that is what this
vector is doing. This common vector over here has multiple things embedded in it. It it it
embeds the the complete understanding of the inputs. The understanding of the input
sentence is very very nicely packaged in that vector. Okay, that's the idea of having that vector over there. Now more
technically speaking, this vector is of a one of the other advantages having this vector is this vector is of a fixed
size output. Your input can have how many ever words you want. Your output can have as many
words as it wants. But what you're doing with this encoder over here is you're taking the input of how many ever words
and then you're always converting that into a fixed size vector, right? And the fixed size vector can be
let's say you know you know thousand dimensions vector at all moments this input is getting converted into a
thousand dimension vector. So that way you'll always know that you can always map any size input into
thousand dimension vector. The understanding of all of that can be always converted into thousand dimension
vector. Then the input into this model is always this thousand dimension vector
and then you're going to predict one word after the other from. So that way your decoder is also decoder also has a
very very standardized input and your encoder always has a standardized output in terms of size at least that's the
other advantage of it. So I think so far we all understand the fact that this encoded
vectorzed representation is exactly what we're sort of trying to accomplish. Right? So all the models that you may
have uh heard of so far, right? Any kind of um large language
models that you may have heard of so far, they all are broadly they all broadly follow the same concept, right?
They all follow the same exact concept. Uh this is broadly the structure of those, right?
There's an input, there is an encoder, you're generating features or embeddings.
Those embeddings are passed into the decoder. The decoder is generating the outputs. Just one more point here. I
mean just for completion sake, there is also outputs that are being passed as inputs here. What do I mean by that?
Outputs from the previous time step are being passed as inputs here. But fundamentally, it's the same. So
nx is n times there multiple n encoder blocks n decoder blocks. So there kind of stacked one on top of the other.
That's what you mean by n. What is below output with red font? Okay. So yeah it's it's the same thing. I mean okay. So if
you take for example a sentence, right? So let's take the input sentence. Start of the sentence. How are you?
End of the sentence. You pass this as an input. You've created this into a feature, a vector. Now, this vector is
being passed into the decoder along with this. There's a first output that is going to be passed. This is the output
final output. What's the what's the first word going to be? First token here going to be
what's the first token going to be? Start of the sentence SOS. So, SOS is going to be the first token. So I pass
SOS over here along with this vector V. This vector V and SOS is going to go in. And what is it going to predict? This is
going to predict in Hindi, right? So it's going to predict the first word as up
H. Now the first one is going to be up. Okay. Then now I come to time the next time step. So what's going to be my
input now? What do I pass here? So I'm going to pass these two as an input now. SOS and up and the vector V. SOS and up
and the same vector V will be passed into the decoder and I'm saying predict the next word.
So I'll predict the next word. What's the next word that it's going to predict? K. Right? So now now I have my
next word. So the third time step I'm going to pass these three as the input along with the same vector V as the
input. Right? Then it's probably going to predict the word HO. Again my Hindi is not the best. So, so don't don't
trust me on this, right? So, that so you want to pres so you want to predict the word ho for the next time step. What do
you do? You pass ho as the input. You will continue to do this until what? Unt until what time? You're hoping that
it stops here or maybe it might predict a question mark. So, I probably will put a question mark here and then I take the
question mark here and then I might simply put a question mark over here and pass it back and this might probably
predict end of sentence. So I'll keep predicting until that particular point. So I'll keep predicting continuously
until I hit end of SQL. Now remember one thing everyone, it's not always only going to predict one word.
When a neural network predicts words, it'll always predict it with a probability distribution. It's never
going to be one word. It'll predict words and probabilities. The last layer is going to be here a soft max.
There's going to be a soft max here. So you're going to be predicting words with probabilities.
So you're never only going to predict one word. It'll be that word plus its prob and along with its probability. So
you always pick the one that has the highest probability over there. So it's not always just predicting one word.
Does that make sense? And in in this sentence, end of sentence will probably have 0.95, but you'll also
have the other words with lower probabilities over here. Your output generally is always going to be a
probability distribution at every time step. It might not always predict the same word.
It'll predict the word along with its probability. So you will have to pick the one that has the highest
probability. The reason why I have put up over here would have been with the highest. That's
why I put the word up here. But if you have a bad model, it might get a very very bad score over there and thereby
this word would probably be incorrect. It might end up predicting it something incorrectly then in that case. Now let's
go one step further. U let's talk about the attention architecture here itself. Right? Right? So the attention component
here itself uh which is what makes this model very very good. Right? So if you actually
look at this detail here, if you now go into the details of this, you see something called as multi head
attention. Right? There are other other feed forward neural networks and everything is like pretty simple
straightforward stuff. But the thing that I want us to understand the most is this concept called multi head
attention. Now we need to understand what this multi-head attention really is. I think
that's where all of the magic really lies. The concept of attention. We all have learned about word embeddings. What
are the different algorithms you may have learned? You would have learned skip gram sibo or you would have also
learned uh word toe. They're all basically algorithms. But the point is when you have a sentence
right W1, W2, W3, W4, W5, you have five words. What you're saying is given a particular sentence,
you can take a particular word, right? You can take a particular word and then you can kind of try to predict
that particular word given its neighbors, right? Given the neighbors, you can
predict given the context, a window of five words, 10 words, whatever that word is, you can predict the word of choice
over here. And we said in the process of predicting that particular word, you will end up
building some understanding of that particular word in the presence of the other words and that hidden vector
becomes your word embedding. So you are essentially saying look if I were to represent all of these words in very
large dimensions then words with the same kind of theme will always come together. A good example is for example
if you take the word milligram kilogram uh and stuff like that they'll all probably come together because they're
all measures. In one dimension they might all be together. In another dimension
they might be slightly far away from each other because they might be slightly far away from each other
because kilogram and milligram are also measures. MIG is small, kilogram is large. So they
also might be far away from each other in a slightly different direction. In this direction they might be close. In
this direction they might be far away from each other. So again the point is you're trying to represent each of these
different words in a very very large dimension. That is what the concept of an embedding really is.
What attention tries to do is attention tries to take this concept of word embeddings a notch further
right um like what what does it exactly do? So if you take any piece of text right
let's take a Wikipedia do document right let's take the Wikipedia article of this let's take any of this article forget
about the images and everything for the moment but let's just let's just take this raw text that you see here the
September 11 attacks commonly known as 911 where four coordinate Islamist terrorist suicide attacks carried out by
al-Qaeda blah blah blah so you see you see the complete piece of text Yeah. Um, ring leader Mohammeda and American
Airlines flight 11 into the north tower of the world trade center you know complex in lower Manhattan at 846.
So you have all the all the piece of information that you see here. Now the thing is when you look at let's say the
concept of the word in this case let's say September 11 when somebody says the word September 11
you know that this reference to the word September 11 is the same as 911 the word September 11 is typically being
spoken about in the context of the word 911 as well now this this reference to the word
September 11 as 911 was mentioned mentioned at the beginning of the sentence, you know, somewhere else in
the in the sentence as well in the in the article as well. And when you look at this sentence itself, right, the
September 11 attacks killed 2977 people making it the deadliest terrorist attack in history. So if you
take each of these words, let's actually copy the sentence. If you take for example this particular sentence here,
you go one word by one, you know, word by word over here. The September 11 attacks killed 2977 people
making it the deadliest terrorist attack in history. So if you just take the sentence like this the interesting part
of a sentence like this is that if you take the word people this word people has some kind of is
qualified by this number 2977. This word people in the context of this sentence is also qualified by the word
killed. Right? And the word killed is sort of qualified by the word attack. So the word people might have a certain
meaning in regular English. But in this particular sentence, the
word people is qualified by a bunch of other information that has been spoken about before this. So the word people in
this sentence might mean something that is slightly different. So the way that you need to attend to the word people in
this sentence has to be adjusted a little. The way when the way when you read a particular sentence when you look
at a particular word you don't look at it dictionary meaning you look at the meaning of that particular word in
context of the words that have been spoken about earlier. So you need to sort of adjust your understanding of
that particular word to the words that have either been spoken about or the concepts of the words that have been
spoken about before it. It's not necessarily only sentiment right? It's not necessarily the word sentiment. Give
you another example. You take the word um you take the word for example um mole right a mole will have very very
different meaning in different concepts the word mole can refer to in chemistry can refer to
6.023 023 into 10 ^ of 23 which is the avagadro's number basically a mole can refer to those many number of atoms or I
would assume that's what it is the word mole in the context of let's say crime or let's say in the context of let's say
judiciary and crime and and that kind of stuff can refer to somebody who's a spy right a mole in the system can also be
referred to as a spy a mole can also be referred to in the context of let's say physiology. The word mole can also be
referred to some kind of a thing on your body, right? It can also be referred to something that's from a mole on your
body as well. Point again being that the word mole could have very different meanings. The word mole can also be an
animal. Yeah, absolutely. Yes, can also be an animal. So it depends very very much on the sentence that it is being
spoken about in right again you need to attend to the word mole very differently in this
particular context within this particular sentence earlier when you learned let's say word
embeddings I mean again I'm not uh we should not be fitting into our own plate right I mean this is a lot of work hard
work that has gone in and I'm talking about how we've improved upon on the work that we've done so far, right?
We've started from the time where you did not even have words being represented as numbers, right? Words
were simply just being represented as numbers using these large wide matrices. Remember document term matrix
where every row simply just has presence or absence of that particular word in that sentence, right? It was a sparse
wide matrix. Words were being represented as very very large matrices as sparse vectors.
From there we went into something called as word embeddings. Word embeddings where you know you pass large pieces of
text into a model and then you try to learn these dependencies through um you know by building that particular model
itself. Um now the problem or one of the drawbacks of word embeddings was that word embeddings learned global
dependency. If you look at glove as the model those are global vectors they were learning global dependencies. So the
mole the word mole would have all these meanings might also somehow be represented using word
embedding. But what it might not represent is but what it might not adjust itself to is if
let's say I have a sentence which says I have or rather this solution has a mole
of let's say I don't know u calcium in it. I'm just making it up here. I have no
clue that that's even a valid sentence. completely forgot my chemistry sessions from school, but I'm just making it up.
So, if you take a take for example a sentence like this, the solution has a mole of calcium. Now, the word mole here
might be referring to if I were to just simply go with word embeddings, it would give me a vector for sure, but this
vector is a generic learning generic representation of the word mole from everything that we may have learned from
a large corpus of data. But in this particular sentence, the word mole, remember, is referring to calcium, is
referring to solution. So I'm probably referring to the word mole. Mole in the context of chemistry. So I would want to
adjust this particular word embedding to a slightly different version of the same embedding. I would want to maybe add
some numbers, remove some numbers. Basically, slightly adjust it to a different representation of the same
word. That is what we mean by attention. So, I would like to attend to this particular
word and its embedding in the context of or in the presence of its surrounding words. So the the way we need to attend
to this particular word in its surroundings will have to slightly adjust slightly change. Now how do you
make that change? How do you make this adjustment? That is what we will discuss on the
other side where we discuss about something called as the attention mechanism itself. Who does the job of
attention? I mean adjustment. There is a model for it. Right? That attention is exactly being
hap that's exactly happening inside this. Whatever you see here, that's exactly
what's happening. You pass raw embeddings and those attentions are computed inside this. Those adjustments
are happening inside the model itself. But we'll exactly understand how that works. The model also doesn't predict
the context. It understands the exact context. It does a fantastic job of extracting that context and adjusting
itself. adjusting that particular word to its context. As per the subject, the word meaning will output the word
meaning will change the the numeric representation of that word will get adjusted. Okay, we broadly understand
the idea of me of attention. Now let's get into the the actual math of it, right? Like how does attention really
work? The underlying core math of it itself. Um I'm going to switch screens because I
just want to switch to one of the um you know one of the topics where Yeah. So I think uh this is your um you
know the encoder decoder architecture. I think we broadly spoke about this and input the encoder decoder with the
output sequence um and then that generates the probabilities as the output. Um so that's pretty
straightforward. Let's actually take a sentence like this. Let me actually write it here.
It's easier for me to All right. Let's take a sentence. Any sentence for that matter. Um let's take
um let's take a sentence like this, right? So if you take for example a
sentence like this um as I said the
obvious way to go forward uh if you were to deal with this with RNNs would have been you would use any of the embedding
models uh you would have created let's say the embedding vectors out of this or train custom embeddings you know from uh
from each of these and then use that to take it forward that would have been the most obvious way uh you know to go
forward. However, in the case of attention, how are we going to do it? We're going to do it slightly
differently. So, let's take each of these words. Okay? So, I'm going to use x1, x2, x3,
4, x5, x6. These are all different words, right? The convenience that I'm assuming here
and this is um a convenient convenient lie rather is I'm assuming that this sentence is fundamentally
getting broken down by words which may not necessarily be true at all times right so when you talk about
breaking a particular word down you would not you would always try to tokenize it meaning break it down to
some of its root forms like for example the word amazing would have become amaze plus ing would have actually been two
tokens instead of one token um but I'm just for the explanation sake I'm I'm assuming this to be like a very simple
um I'm assuming this to be sort of a very simplistic uh um representation here in
this case right so let's take for example uh day as an example if you take the word day
x5 for that matter as I said the word day would have had a default default embedding right whatever that embedding
is let's consider X5 X7 these are the default embedding these are the right these are the default embedding right so
this is understanding attention these would have been the default embeddings and these are embeddings that you would
have generated through let's say any of your traditional embedding algorithm or you can actually train those algorithm
train those embeddings as well doesn't matter the question is how do you adjust this
particular particular word day. I'm taking day as an example here. If you take for example this particular word
day, X5, this word day is not in its best shape here. We would probably have to adjust the word day to ensure that it
also gets adjusted to the word embedding amazing here because the word day is being qualified by the word amazing or
the word amazing is qualifying the word day. Amazing is the adjective. Day is the noun itself. So the this particular
noun is being qualified by the by the word amazing. Um similarly the word they is also uh dependent somehow on Dave
because Dave has you know Dave is the one that has actually had an amazing day. So the word
day is also dependent on the word Dave in some in some shape or form. Um and so on and so forth. You can think of how
some of the other words are also dependent on the others. So how do you adjust these embeddings?
So to adjust these embeddings right, one of the most simplest ways to do this is we say, you know, you could essentially
take these words um and you can think of
trying to somehow multiply these embeddings that you see here, right? somehow and take these embeddings. So
for example, if you take X5, I'm going to create something called as Y5, which is an adjusted
embedding of X5 or of the word day. And I'm going to say Wi-Fi is somehow a numeric representation of W1, X1. So it
is a combination of the embeddings of all the other words. W2 X2
W3 X3 plus all the way until W8 X8. What am I saying here? All that I'm saying look this word wifi or this word day
has a new embedding representation. This embedding representation is not the same as the older representation.
This representation is a linear combination of all the other embeddings that I have here. Meaning in some shape
if let's say the word day is heavily dependent on the word amazing then X4 or W4
would have been a very very large number and W4 would have contributed to X4. So W4 would contribute more to the word X5
and thereby the summation would have been much larger in this particular case. Right? W1 would probably also be
very large. Maybe the word W2, W3 might be actually much smaller. W8 or W7 might also be very very small because these
words may not necessarily contribute a lot to the word over here which is W5 or X5.
My point being that the new adjusted embedding could be created as a linear combination of the original embeddings
itself. So this way you could potentially get all of your new embeddings,
right? You could potentially get all of your new embeddings as a linear combination of your original embedding.
But the question is from where will you get the W's? So your final embeddings are over here
are going to be instead of x1, x2, x3, x4, the final embeddings that you would probably be working with is probably
going to be y1, y2, y3, y4, and so on all the way until y8. These are going to be the final um embeddings that you'll
probably be working with. But the question is, how will you get these W's? Where will you get these W's from? Who's
going to give us these W's? Where are you going to get these W's from? And that's what we'll discuss. Again, it's
those W's that will tell us in what form or what combination do you need to combine these existing embeddings of
these sentences to create a better embedding or a better representation of that particular word in this context.
Sharam, it can be any of your existing embedding models. It can be a word to model. It can be a skipgram gram model.
It can be any of the models or it can be a fresh embedding itself right you can just train a fresh embedding itself for
all you know every new embedding now which is Y1 Y2 Y3 all the way until Y8
in this in this context these W's are to be identified and we need to understand where these W's will come from these W's
are nothing but the relevance of a particular word in the context of this particular setup in this in this
sentence it's the context of the word that is associated with the others. But what I what you need to understand is
these words or these weights are specifically these weights over here
that you see you know when you create Wi-Fi these weights are specifically for Wi-Fi. Similarly, you will have other
weights for X1. You will have similar weights for X2, similar weights for X3, similar weights for X4, for X6, X1 and
X8 respectively. My point is you will have different sets of weights that will combine for each of these individual
vectors and thereby creating the final output. That's one thing that you need to visualize. That's one thing you need
to understand. Now, let's go one step further. So, how do you get these weights? Where will you get these
weights from? So as I said these weights are essentially
I mean it's it's not very straightforward. Um there are some very simplistic ways of thinking about it.
But I'll tell you the most um for a lack of a better word I'll tell you the most uh
um you know the the actual technical way of getting to the final weights itself. So the way we will get to these final
weights is remember these weights have to be trained right there is no rule of thumb you know you cannot just randomly
get to these weights right away. So to get to these weights you will of course have to
go step by step. Um there are uh other sub weights that are sort of created. Let me give you a simple example here.
Um so to get to let's again let's stick with uh any of these let's stick with Wi-Fi
or whatever. As I said we will have different sets of weights that we'll introduce to compute W1 W2 W3 all the
way up to W8. But how will you compute them? If you take for example let so we will be
introducing two sets of new um two sets of new matrices called the query vector or the query
matrix Q and the key m key vectors. Now what is the query and the key vector
respectively? So for example, if you take the query vector, what does the query vector here mean? If
you remember, what did I tell you? The word day, right, is being qualified by this particular word, amazing. Amazing
is the adjective. Day is the noun. So one of the ways to think about it is
okay, what are the words that are qualifying the word day? You somehow need to find what words in this
particular sentence are qualifying the word day. Once you know what those words are, then you can use those the
respective vectors and then some somehow combine it. But how will you know what those what those words really are? You
would of course not know it. We will we will have to learn those using these vectors. So I'm going to be introducing
something called as QI which is nothing but a query vector which is going to try and query for okay what are the words
that are qualifying the word day what are the words that have the dependency of this particular word day. So that
will be WQ multiplied by X5. Right? This is specifically for the
word day. Right? So Q5 is WQ * X5. That's the first one. Then you would also have a key vector
for each of the other words. So K1, K2, K3
and so on and so forth. Now what are these key vectors? These key vectors are simply nothing but they are
essentially going to carry the value rather these key vectors are going to be the actual values of these
u inputs itself. So k1 is simply nothing but w k * x1.
Similarly, K2 is W K * X2 all the way up until K8 is W K * X8. So now what's going to happen is
we are now going to try and multiply the query and the key vectors over here. So the way for you to understand this is
like a large matrix. So you have the key vectors K1, K2, K3 all the way until K8 and then
you have the query vectors Q1, Q2, Q3, Q8. So you have Q5 here. So you're
essentially going to take, you know, the values of Q5 multiplied with K1. You're going to
multiply Q5 with multiplied with K1. Q5 multiplied with K2. Q5 multiplied with K3 all the way until Q8 multiplied with
K Q5 multiplied with Q K8. So the query and the key vectors essentially going to multiply themselves and wherever you see
that the value is very large meaning the inner product the multiplication of both of these vectors if you think is very
large that's an indication to say that look this particular vector right this particular vector which is
word pi this input vector or rather sorry not wi this input vector xi has a large so for example if this
number is very large and this number is very large. It's a way to tell you that hey you know what looks like X5 has a
lot of dependence on qk2 or rather w2 and also on w8. It's a way to identify that the inner product is very large.
It's a way to identify that these two vectors are have a lot of dependency between each other. That's what we are
sort of trying to get at. In a way we're trying to find dependency. In a way, we're trying to find some level of
overlap between both of these, right? X5 and or rather each of the words so to say. I just want to keep the notation
the same here. I don't want to confuse you all. X1, X2, X3 all the way until X8 and so on and so forth. So, in a way,
you're sort of trying to combine the overlap. Okay? So, I'll simplify this. Take any two vectors. Take any two
words. Okay. Let's let's actually go back here. Let me explain the concept and then we'll get into the math a
little. The concept here is the intuition here is if you take any two words these two words if they are
similar the concept here is that if you take for example any two words right like for example in this case the word
day and let's say the word amazing. If they have something in common between each other, when you combine the two of
them or when you take an inner product of both of these vectors, it would typically yield a large value. If there
is anything in common between both of these, the dotproduct between any of these two vectors should ideally yield a
large value. Do we agree with that? If there are two vectors which are similar, if you
multiply both the vectors, any bit of commonality between the two would yield a large value. So in theory that is what
we're trying to understand. That is what we are trying to do here. We are facilitating that multiplication.
We are creating one query vector which is the vector that we are trying to find similarity for.
This is the word XY that we're trying to find similarity for. Right? And then there are these other candidate vectors
which are other words which we're trying to find similarity against. Right? Right? So we are trying to basically
find similarity for X5 with X1, X2, X3 and so on and so forth. We don't want to take the raw vectors themselves because
the raw vectors themselves because again these are vectors that are um coming out of an embedding vector. They you know
you might want to let's say multiply them with other you might want to let's say either shrink these vectors down.
You want to reduce them in terms of dimensions. That is why you're providing these other word you know WQ and WK to
ensure that you're not again this is more for computation as well and to also sort of control
uh to also adjust for what kind of information you want to pass into it and what you don't want to pass into it.
These are tunable parameters. Um so for that reason you just multiply them with another vector just to kind of make sure
that um you're passing the right or you're extracting the right information out of it. Those are like regulators.
Think of these as regulators, right? But fundamentally what you're doing is you're essentially multiplying
each of these vectors against each other to see if there are let's say 10 vectors, eight vectors here. So x5 is
multiplied with all the eight vectors to see where the similarity is going to be the highest.
If the inner product if the dot product across both of these is very very high anywhere right if Q5 multiplied by K2 is
very large what that means is somewhere Q5 and K2 have some commonality with each other in a way you're trying to
compute some kind of a correlation not exactly correlation some kind of a correlation you're trying to understand
so you have the query and then you have the key vectors the query vectors are the vectors that
you're trying to find similarity for. So these are whatever you see at the top. These are query vectors. These are your
key vectors and you're trying to find similarity against each other. Then once this multiplication is done,
once this multiplication is done, you're of course going to perform this multiplication against
every uh you know every word against every uh other every query vector against every
key vector. Okay. And after that multiplication is done you're going to get by the way this value that you see
here that multiplication we will simply refer to it as zed or the new vector we simply represent it with a small z here.
Um so Z1 is simply going to be um in this particular context Z1 is simply going to be Q you know K1
multiplied by Q5 or in a way simply put Z1 now this is K1 multiplied by Q5. This is
Z 1. Similarly Z2 is going to be what? K2 * Q5. These are
essentially these values. This is Z 1 PI, Zed 2 Pi, Zed 3 Pi, Zed 8 Pi. So to say that's essentially what these values
really are over here. And then if you then then how do you get the dependency out of this? Then how do you get the
dependency out of this? So after what you get all the values of zed here after that you simply apply the final weights
are going to be what? So the weights w1, w2, w3 all the way until w8 for the vector phi or for the word phi are going
to be a soft max of z5, zed 2 all the way up until
z 85. So you're essentially simply normalizing all of these values that you see here. You're simply normalizing all
of these values here. How many values will you have here by the way? You'll have a total of eight values. So that
will give you all the vectors, all the weight. So when you do a soft max, what are you going to get? So you will
normalize the values. You're not picking the values with high probability. You're normalizing it. Argax will pick the one
with the high probability. This will convert everything into probabilities or normalize the whole thing to one. It'll
basically normalize the values to one. Does that make sense? Because these are simply uh dot productduct. So it can
vary from vary from negative infinity to positive infinity. So this will simply give you values like you know w1 all the
way up until w8 for the vector 5 will be something like 0.18 2 or rather sorry 02
06 01 and so on and so forth. You'll have values like that. That way you'll get
the weights. But this is specifically for five. No, no, there's no average. This
is for W5. Sorry, for X5. So what is the final YI? Whatever weights that you've got here.
W1 * X1 plus W2 * X2 plus W3 * X3 plus all the way until W8 * X8. This is the new vector Wi-Fi. Remember
these weights over here. are these weights that you got from them. Is this clear everyone? This is
only for Wi-Fi my friends. U artificial intelligence is transforming the human. So if you look at the word artificial
intelligence look what happened artificial intelligence is transforming the has actually been broken down into
um artificial is broken down into art and eicial intelligence is transforming. This is how it's got tokenized. By the
way, there's a there's an algorithm that is used to tokenize it. Then there is a token embedding. Basically, just a
number. Every every token has its own embedding, right? So, if you think of it, there is some kind of an embedding
that is created over here. 768 long vector embedding converts tokens into semantically meaningful numeric
representation. How does it come? It comes from any of your word to models or it can be a
simple numeric representation. What's the size of the input vector here and everyone? This embedding is of what
size? Each word is of the size 786 or 768 each word. So this is x1, x2, x3, x4, x5. Each vector is of of size 768
here. After that, there is something called positional encoding. We'll come to positional encoding in a minute.
Let's not let's not worry about positional encoding. It'll be too much for us to worry about for the moment.
Let's just let's just ignore personal encoding for the moment. It's just a way for it to embed the size of the personal
encoding. Now let's get into some of the detail here. So the QKV computation
um so let's go one by one. By the way, here are the attention weights that are coming out. Okay, let's
let's let's understand one by one. You forget about the value over here. You forget about this
particular value here for a minute. Just focus on the Q and the K vectors. Right? The Q and the K vectors.
The query and the key vectors over here each are of the size. Let's take any of these.
Let's take one of these. Just a second. I'm trying to get into a little bit of detail. Yeah,
this is residual. That's fine. Yeah. So, here's the connect, you know, here's the computation that you see here. Um,
what is exactly happening in this particular process is if you see this, this word art has its own embedding.
This is getting multiplied by as you see here. Yeah. This is getting multiplied by the
query vector. Similarly, this query vector over here is
available. Um, the problem with this is that it doesn't hold for a second. This query vector is getting created
here. So, E11. So, whatever you are observing here is Q * K dot V. But you forget about the V for a second. Q do. K
is what is getting computed here. You know, whatever you're observing here, Q.K is what is getting computed here in
this particular setup in this particular in this particular step. Um, and that is being combined. So
whatever you get out of the key is multiplied by what is coming out of the query. Both of these are getting
multiplied and then the both that dot product is being computed here. So this is zed for you. This is the values in
terms of zed for you. Whatever you're seeing and these zeds you're applying a soft max over here
these zeds are essentially going through soft max over here right that soft max. So if you see here
if you look at that computation over here Q do. Kranspose or rather Q do. K your query and the key
vectors are both getting multiplied and then there's a soft max being applied around it. It is divided by root of you
know under root DK and there's some mathematical nuance over there that is for smoothing and and and and those kind
of purposes. Don't worry about that for the moment. Again, minor detail, but as you can see, the output that you're
getting out of here is a softmax output that ranges from minus1 to minus1 to + one. It'll typically be only uh 0 to
one. In most cases, it just sum up to one. As you can see here, all of these values are simply going to sum up to
one. That's it. So that those are your attention weights. These are your W1,
W2, W3, W4, W5, and so on and so forth. You're taking the output from the key, output from the query, multiplying both
of it and then you're trying to get to a certain number. Now what's interesting here
is what is interesting here is as you see there is a part of scaling that h sorry
masking that happens here. What do you mean by masking? See when you're computing something for the word for
this particular word right? If you take for example the word If you take for example the word art.
The word art cannot have cannot be dependent on any of its future words. Art cannot be dependent on
facial intelligence is transforming the because art is being spoken first. Right? Similarly here if you take the
word artificial artificial is not dependent on any of the other future spoken words.
The vice versa is possible. If you take for example any of this, if you take the word 'the', this word 'the' can be
dependent on any of its past words. Which is why all of the upper triangular matrix that you see here, they've all
been forced to sort of a zero. If you see here, they've all been forced to zeros. those do not
contribute to your attention at all. This upper triangular matrix over here that will never contribute to your um
you know to your uh you know to your attention values itself. Exactly. Most of the words dependent on the past word
they don't depend on the future words which is why you simply just get rid of those values. Um that's the reason um
they are not they will actually be made negative infinity. The reason Ashish they're made negative infinities is
because when you apply softmax they'll be become they'll be made zeros. So technically in the process they're
actually made negative infinity. You forced them to negative infinity. That way you can ensure that the soft max
will will push it to um you know zeros because you again want the summation to become one right so soft max of negative
infinity is zero but anyway that's essentially the output guys that's how that's how this thing works this my
friends is the attention part of it that's the first step the other thing that I want
to talk about is if you look at the by the way The QV weight matrices are all of If you look at here, the QV
computation that you see here, the QV computation that you see here, if you look at the we'll come to the V matrix
in a minute, but if you look at the Q and the K matrices, they're each of they are square matrices 768 by 768.
Why 768? See, because the input matrix is of size 6x 768.
Each word is 768 vectors in long size and long. So when you have a query matrix which is of size 768 by 768 when
you multiply you will get a when you take for example a do uh you know when when you multiply
this 6x 768 with 768 by 768 um you will of course get a matrix that is going to look like this.
um you know uh you you'll of course get each of this is essentially going to become a size uh that big um you know 6
by 768 again in terms of rows um and and all of those are essentially stacked against each other and then you simply
summed up um you know multiply you of course take a dot product against each of those um that's how you get um you
know two parts that that's how you get to summation but again don't worry about the the underlying too much detail. So
this by the way is first head. This is one computation of QV Q and K. Similarly
the same thing will happen across 12 different such processes. Whatever we are seeing here the same thing will
happen multi computing attention will happen across multiple heads. Heads meaning simply multiple blocks.
The same thing will happen across multiple multiple blocks which is why this is referred to as one this is
referred to as self attention. Why is it called self attention? Because you are computing the attention of a particular
word with its own self. Then multi- head self attention because you're computing the self attention
across multiple you know this process is repeated multiple times. This process of computing this attention
is repeated multiple times. Um and you're also applying masking 12 is just a hyperparameter one. It's like uh how
you have um in a VG16 why do you decide to have 16 layers? U same thing here it's it's just that it's
it's a parameter you can change it and then you have of course multiple heads. So which is called multi
head self attention. But more importantly um this is also there's also some level
of masking that's happening here. So it in a way it's also referred to as masked multi head self attention or self
attention with masking. Uh it's also you know the term masking is also used because this forcing the upper
triangular matrix to zero that's referred to as masking. So you just mask the values that are going to have any
future dependence. So which is why it's separated. uh masked multi head self attention or uh self attention with
multi head self attention with masking. Then comes the last part which is the value. So we looked at the
query matrix, we looked at the key matrix. There's also one more vector matrix over here called the value matrix
which is then multiplied with the output of the attention. Whatever output comes out of the attention is further
multiplied with this value matrix. Whatever value matrix value matrix is also a 768 by 768 premium sized matrix
again whatever you see here um that value matrix is again further multiplied with the uh if you look at this
um again the same same matrix is sort of multiplied here again. So when we actually go back here, the
only difference is that all of the comput sorry all of the computation remains exactly the same.
The only thing is here after all of this computation is done. this yf5 that you see here which is nothing but summation
of uh w i i xj
right this is also multiplied by vj which is nothing but a value matrix right you just don't leave it here you
also multiply it with another matrix over here another value matrix again another form of a regulator it's another
another regulator here in this case um so that's where the value vectors sort of Yeah, it's I think to 2 million now.
So maybe I might have gotten it wrong. Gummini Pro is probably 2 million like what Ashish is saying. So that's the
query key and value vectors. So once the query key and value vectors as you can see the output whatever you get out of
query and key. So you multiply query and uh query and key. This is a simplified computation so it might look simplistic
but whatever you get out of query and key you further multiply it with the value. Whatever you get as an output as
you see here the output each of it is is size 64 in terms of you know size that is sent out for the subsequent steps.
What are the subsequent steps here? Um just the way we saw by the way this is for one block. Similarly there could be
other um blocks as well over here. There are other aspects that we will talk about right? So there is um things like
for example from here there's a multi simple multi-layer perceptron very very simple block that you see here
just uh let's go in here the a regular MLP hope I can zoom into this
so whatever u the point is whatever output that you get for each of so this is the so the only thing that we've been
able to accomplish you started with a 768 long vector After the attention you get a 768 long vector
again nothing changes started with 768 ended with 768. So what happened here in the process you've adjusted the vectors
for each of these individual input tokens. That is what has happened. So this whole block that you see here this
whole drama is just to adjust the input vectors of each of the input token into a newer representation.
That is what has happened here. After that we then put it through a bunch of other
things. Right? So the first step here is you put it through a simple multi-layer perceptron.
Um a multi-layer perceptron over here and and this multi-layer perceptron is like uh with dropout you can have
multiple again these are think of consider these as multiple inputs. Each input can be passed into a regular
neural network multi-layer perceptron regular neural network with a regular feed forward neural network. Um there
are residual connections here. Um if you've learned reset you would know what what that means but simple residual
connections. Again not very complex simple residual connections. Um and then last but not the least you get the
output from here. Again the output also as you can see here each of the outputs is again 768 long vector right um then
all that you do is whatever outputs that you get here whatever softmax outputs that you're getting here by the go back
here one output that you're getting here you will average that across 11 other transformer blocks so whatever block
that you're seeing here whatever block that is getting highlighted in blue there are 11 other blocks because there
are total of 12 blocks all of those blocks are getting summed up here. So if you just look at uh yeah so this
whatever output that you're seeing here or the for the word for the word field um you're essentially summing up or
averaging the output across all the other transformer blocks and then passing it into a simple feed forward
neural network or rather into a simple softmax layer and predicting what the word would be next whatever is the word
in this particular case. The artificial intelligence is transforming the human. Human is the input word and the output
is human field. Should it be field? Should it be the output but the word with the largest is
way. So the output here is going to be taken as way. In this particular case the word the output word here is the
human way. Transforming the human way. That is the word with the highest uh uh probability. So that might be picked up
over here. This is a transformer encode or rather this is a simple transformer model. Whatever you're looking at over
here, this is the process of converting an input into an encoded representation that is then
further passed into you can either use that to pass it into a softmax and generate an output or you could pass it
into a decoder and generate further output. We'll talk about the rest later. What we saw is you passed an input,
passed it into a multi head attention, passed it into a feed forward neural network and you got a numeric
representation as an output. You got a vector as an output from the encoder block. You got this vector as an output.
If you know how the encoder works, then it's the same piece of stuff on the other side as well. What we saw here is
a far more simplified representation of it. Of course, what we are seeing here is a far more simplified transformer
block. But what you could potentially do is you can take that output you can take that particular output,
right? And then you combine that with your output embedding whatever words that have come out so far. You compute
the multi head attention for those. You merge all of them. You concatenate all of them over here through attention.
So the whatever input you're getting plus the word that you're getting output over here plus the output word you
combine all of it and then pass it into a feed forward neural network into a softmax layer generate the output simple
the add and normalize is simply these are add and normalize is layer you know is batch normalization layer
normalization batch normalization that's what's happening this is sort of layer normalization in a way and the
outputs of that particular layer are simply getting normalized that's what's happening here nothing
Right? This is this is very simple and these connections that you see here they are residual connection. So broadly
speaking whatever you are seeing here so this part so this part is whatever is highlighted
is your encoder. Whatever is highlighted is encoder. This part
is sort of the decoder part. Now how you want to use the decoder is totally up to you. Either you can just have an
existing decoder or you can because you are also going to pass new word as inputs over here. You can use that for
decoder especially if you're doing it for sequence to sequence sort of output. Where is the parallelizing happening in
transformer block? So, so this QV so what we have completely eliminated here is we are not treating
language anymore as a sequence of words. I mean it is still a sequence of words but you're not treating it like a
function of time. We're not saying hey you know what you need to treat the first word then go to the second word
then go to the third word then go to the fourth word. You completely took that concept of recurrence out of the
equation. All that you're doing here with attention is you're treating every word all at the same time. Exactly. All
the words are treated at the same time. And because all the words are treated at the same time,
everything can be simply matrix uh everything is matrix multiplication. Now does that make sense? So you suddenly
started to treat sequence of words as just a simple matrix which is fantastic because you kind of took the whole idea
of sequence context and everything and you somehow packed it into this idea of uh attention
and because of the beauty of attention you took that concept of recurrence out of the equation and now all that you're
doing is is just using attention to uh you know to capture all that uh you know sequential dependencies and you're sort
of using masking a little smartly here because the way you're using masking you're sort of also capturing that
sequence in some form or shape because of this concept of masking you're able to capture that dependency
very very well >> just a quick info guys intellipad offers generative AI certification course in
collaboration with IHUB IT ruri this course is specially designed for AI enthusiast who want to prepare and excel
in the field of generative AI I through this course you will master geni skills like foundation model, large language
models, transformers, prompt engineering, diffusion models and much more from top industry experts. With
this course, we have already helped thousands of professional and successful career transition. You can check out
their testimonials on our achievers channel whose link is given in the description below. Without a doubt, this
course can set your careers to new height. So visit the course page link given below in the description and take
a first step toward career growth in the field of generative AI. We're going to discuss the applications of the
transformer architecture. Right? So, as I said, there are many many many applications of the transformer
architecture. Um, one of the most popular ones these days are the GPT models. Of course, before the GPT models
came up, right, the GPT family of models came came up. There was another family of uh transformer architectures which
was actually fairly popular. Uh so Google had launched this uh model called the bird model which kind of became
super popular. Um yeah let me pull this up. Uh I'm trying to find the right u their homepage but
okay let's let's look at this. Yeah this is the paper. So the BERT pre-training of deep birectional transformers for
language understanding. This was exactly the model that was uh that was launched by by Google. Um again a a very very
revolutionary paper at that time. Um the idea of the BERT model was exactly uh you know to improve the
the language understanding capabilities of uh from a natural language processing standpoint. And of course transformer
models are really good. And uh this this architecture actually was doing very very well. Right. So the BERT model was
actually one of the first models that had uh that had really shaken things up a little. Uh
it's very similar to your uh it's it's very similar to your traditional encoder decoder architecture that we just
learned. Let me actually let's actually go to the paper itself. Uh this was published as you can see in 2019.
The original model was published in 2017 whereas this model was published in 2019. Um and then uh yeah the BERT model
is actually very similar u to that of your regular uh any of your other transformer models that you that that we
may have discussed so far. It's not very different. Uh it's just that this particular model was used for a bunch of
others you know specific tasks um you know so to say. So I I'm not going to go into too much detail here but I'm just
going to you know talk about uh specific concepts around uh the BERT model itself right so we'll actually look at how a
BERT model can be implemented let's actually start with a very very simple uh let let's start with how you how one
could use let's say some of these transformer models interestingly enough as I'd said a lot of these transformer
models were launched over the last few years, right? So, last two to three years, what we observed is a lot of
transformer models were launched. They were all based on the same concept that we had discussed, right? They're all
based on the same exact architecture, but you know, making some minor changes here and there, different languages. Um
some of them are being used for classification, some of them being used for uh um some of them being used for uh
specific natural language processing tasks, some of them for embedding creations and so on and so forth. Um so
what has happened is as an outcome of this so on one side the model architecture is very popular on the
other side what was also being discussed what's what was also being um you know introduced is a lot of variants of these
transformer models the idea now is for us to go one step further and and to see if we could um combine um all of these
and if you could create like a a simple interface where you could access all of these different models. Um so the the
interface that sort of came out to be one of the most popular interfaces um in this particular case is the is the
hugging face um uh architecture. Let me just let's go here. We'll be taking this one step further. We'll of course be
trying to understand how the transformers library can be accessed using this uh
library called hugging face. Right? So there's a very very popular library called hugging face that sort of came
out to be or rather a a group that had launched a library called the transformers library. The primary
purpose of the transformers library was just to make all of the transformers model available through a very very
simple interface, right? Through a very simple library like interface. Let me actually quickly show you how this looks
like. Uh so here you go. There you go. So here's the transformers library. So so the transformers library as I said is
a is is very much like your scikitlearn. In your skarn, you just have access to all of the traditional machine learning
models. Whereas the transformers library here provides you access to all of these models that we spoke about, right? Be it
BERT or let's say any of those other variants that you may have think of. All of those models are nicely available for
all of us to access. Um, so you could do all of these different tasks for example like natural language processing,
computer vision, audio, any of the other multimodel tasks as well. you could use the transformers uh library to kind of
go through uh each one of those. Um so let's let's exactly understand um how one could access any of these different
models. By the way, uh as I said there are a bunch of different models that are available. So for example, if you want
to access a model called BERT, this BERT model can be accessed through PyTorch can be accessed through the transformers
library via TensorFlow. You can also, you know, access it through flax. Flax is another uh another modeling software.
You could use that as well. So, a bunch of different models can be accessed through all of these different um these
these these different programming languages. The default as you can see here is PyTorch. PyTorch is of course
the default uh model or or a framework that can be used to access um all of these different um um or all of that can
be used to access all of these different models. Let me actually quickly give you a let let me give you a very very quick
introduction to hugging face um and then we can take it from there. Okay. So, let me quickly show you how you could access
all of this. Let me Perfect. So, I'm not I'm not going to get into like a lot of detail here, but I'm just going to
briefly touch upon this. So, what is hugging face? U hugging face is this is uh you know sort of a tricky question uh
because uh hugging face today the modern day setup has a lot of offerings right. Right. So, hugging face is is also a you
know it is a it has libraries. It has a platform where it is also like a community uh it it is like an
open-source repository of models, data sets and so on and so forth. So, it could do like a bunch of different
things. Um here are some examples, right? So the core offerings of uh of hugging face they provide you a
repository of models open source models they provide you a repository of data sets they provide you something called a
spaces which is like for you to simply go and execute stuff right and of course you have documentation libraries you
know and a bunch of different things right so you have a community and and all of that um so what we're going to be
doing is I'll just quickly show you how you could access some of these models, right? Uh hugging face today hosts a
bunch of different models, right? So you have the large language models, your you know your diffusion models, text to
image models, so on and so forth. So all of these base transformer models as well as the more complex large language
models, diffusion models, all of those are completely available for someone to access using the platform called as
hugging face or using this library called as hugging face. Let me quickly show you an example of how one could do
this. Right? So if you like for example here, let me go to the website itself. I don't want to let's go all the way up
here. So if you go to models, right? So if today if you just go to the models tab here right in the models tab you see
all of these different models that are available right so you could simply click on any of these models and you can
pick up and then you can understand okay what is this particular model so this is a reflection llama model let's actually
look for a simple bird model I don't want to over complicate this let's go text classification
okay let's let's look at this model called distal BER and I'm trying to find a BERT model itself. Okay, so there you
go. This is the BERT model. This is the older BERT model. BER large uncased whole world masking fine-tuned squad.
This is essentially saying the BER large model uncased meaning it is trained on uncased data set. Whole world the
complete word is being used as a token masking. So there was some kind of a masking that was taken place and it was
fine-tuned for squad. There's a data set called squad and it was fine-tuned for that particular data set called squad.
Um, so you can actually read about this. You see something called as a model card here. Uh, pre-trained model on English
language using mask language modeling objective. It was introduced in this particular paper. Uh, firstly released
in this repository blah blah blah. So there's all of that information that's available for you. Let's actually take
one example of the BERT model itself. This model has the following configuration. And it's a 24 layer
model. 1024 hidden dimension, 16 attention heads. Um 336 million parameters. So 16 attention heads. What
does that mean? Spoke about multi head attention. Remember multi head attention? We spoke
about multi head attention in our previous session. Exactly. So what we're saying is 16 attention heads are there.
In the example that we may have seen, it would have had uh if you remember that should have had around it would have had
12 attention heads. They would have spoken about 12 attention heads here. However, the Google BERT model has 16
attention heads. That's how they have fine-tuned this particular model too. Okay. So, uh this model should be used
as a question answering model. You could use this particular model for doing any kind of Q&A uh any kind of question
answering um and um you can use this for performing specific uh you know question answering setup given a certain corpus.
You can use this particular model for doing any kind of question answering. So let's let's exactly understand how you
could use any of these pre-trained models. Remember these are all pre-trained models, right? So what you
could do using this model is you could use uh you can pass a paragraph like this. You can ask a question like this
and it would try to answer from this particular paragraph. It would try to answer this particular question from
this particular context. So given a piece of context and a question this model tries to answer that from this
particular model itself. That's essentially how this model works. As you can see here, uh, which name is also
used to describe the Amazon rainforest. Um, and you have like a complete paragraph that has all of the context
here. You can simply say compute. Um, I need to log in here. Perfect. Compute. And it would execute and it
would return an answer saying Amazonia. Which name is also used to describe the Amazon rainforest in English? Amazonia.
The Amazon rainforest. And all of these all of these uh names also known as in English as Amazonia as
you can see here that is the response and it has come back with that particular response over here. So that
quest kind of question answering can be done using a model like this. Remember this model is the same as the one that
was published here. Right? So it is exactly the same sort of a model that was discussed uh earlier as well. This
particular model is the same transformer architecture only that was used but given a piece of context it tries to do
question answering. Instead of generating new content it does question answering that's the only difference. So
there are different tasks that this particular uh bird model has been trained on. Right? There is no left to
right or right to left language models to pre-train BRE. Instead pre-train, we pre-train BERT using two unsupervised
tasks. There are two specific tasks that they've taken to train the BERT model. They they don't just do word prediction.
Instead, they use something slightly different. They don't try to predict the next word. In this in the case of the
BERT model, they've trained it using a slightly different technique. What are the different techniques? Something
called as masked language modeling. The idea of a masked language modeling is given a particular sentence they kind of
mask a particular word it's like fill in the blanks right so you take a sentence you try to make one word you mask one
word and you try to populate that particular word and that that empty word is at random right uh so intuitively it
is reasonable to believe that deep birectional models is strictly more powerful than either a left to right or
right to left model because you birectional your learning flow from left to right as well as right to left.
Unfortunately, standard conditional models can only be trained left to right or right to left. Hence, we train it
using something called as mask language model. Uh how does it work? We simply mask some percentage of the input tokens
at random and then predict those masked tokens. So, it's like fill in the blanks. take a sentence, try to make one
word empty and then you try to predict what that particular word is. Uh and this this procedure is referred to as
masked language modeling. This is exactly how these this particular model has been trained. So in this particular
case around 15% of the tokens have been masked at random given any particular sentence and they try to populate it.
Now this is a way to learn the the relationships between words given any particular sentence. That's how they
learn the relationships. There's another task also that they use called next sentence prediction. Many important
downstream tasks such as question answering and natural language inferencing are based on understanding
the relationship between any two sentences. The other task that they this model has also been trained on is to be
able to predict the next sentence. Um given a bunch of sentences, it also tries to predict that next sentence. So
this is how the BERT model has been trained. It is slightly different from your regular transformer architecture
itself. That's one point that I want you all to understand. But anyways, going back here to the case of uh your uh
BERT, sorry to going back to the case of your uh yeah to the BERT model itself. Um let's quickly understand how you
could use uh firstly the transformers the the hugging face uh interface um and then or rather the hugging face library
and then secondly what we'll also do is we'll try to perform some kind of a quick uh we'll also try to perform like
a quick question answering sort of a setup using the BERT model one of these BERT models.
There you go. Let's go. So uh bag um exclamation pip install transformers.
So you will have to run the pip install transformers. So in this particular example,
all right, here you go. So this is a transformers library. Um so what I'm doing here is just look at
this. So from transformers import pipeline. Pipeline is like a default
default object that is available. U and I'm saying hey in pipeline I want to do something called as a sentiment
analysis. Um and what it's doing is it is fetching this particular fetching a default model
for this pipeline exercise over here and it is trying to perform the classification over here. So it is
trying to do sentiment analysis. Um so as you can see please install a backward compatible TF
car package with pip install TFAS. Let's do that. Let me switch or else let me just go to
PyTorch. This should be able to execute it otherwise. Yeah, there's a ver version
mismatch of the libraries. Yeah. So it's using this model called distl but base uncased fine-tune SST2.
So it is using this particular model. It's going to this particular model. This is the default model that is being
utilized to perform this particular classification. So what am I doing? I'm saying hey pipeline and in this
particular pipeline what we are saying is it's a factory method um in the case of hugging face so it takes two aspects
as inputs it takes one called as a tokenizer and the second called as a model it takes both of these as input I
don't it is not mandatory for me to provide it there is also a default that is available for this which is what is
currently being utilized and then after that it is able to perform a classification itself. So the
classifier takes this particular sentence as an input and it performs the classification. It's saying hey the
classification for this particular model is negative with a score of.99. The experience with the Apple customer care
has been horrible and as you can imagine this is a negative sentiment and it has returned a negative score for this
particular sentence. In this case, what's happening is it is taking this model
uh and performing a simple sentiment analysis exercise on top of this. Now I have not trained the model here. It is
using an existing model. It is using an existing model you that is already available on hugging face. Downloading
that particular model and simply performing a classification using that particular model.
I have not performed supplied any model but as you can see it has defaulted to this particular model. So the default of
this pipeline model is of this pipeline class is this particular model dist base base uncased fine-tuned SST2 English. So
this model is the default model that is available. It has simply performed that classification for us over here. Right
now I'm only doing inferencing here. I'm not doing anything else. I'm just taking the model, taking my sentence and
performing the side classification here. I am not building the model. This is using a pre-trained model. In this
particular case, now I'm going to say going to perform something called as a zeroshot classification.
This is the score of this prediction probability score. Let's do a zeros classification. What do you mean by
zeroshot classification? See, there are multiple ways of performing this particular classification. So to perform
a typical classification, you would take a data set 10,000 observations, pass it into a model, train the model to perform
the binary classification or or threeclass classification or multiclass classification. That's the regular
approach to perform a classification problem. But what you can also do using some of
these models and and that's the beauty of how these models are is you can just take a pass a sentence and you can say
hey look I need to classify this particular sentence into one of these three categories. I I tell it nothing
else. I pass a model I pass a sentence and I say hey take this particular sentence and classify it into one of
these three categories and that's it. That's what it does. It takes this sentence and performs the classification
for me. And in this case it's using Facebook Bart LG uh MNLI which is another model that is being used for
performing this particular classification. Who has decided what model to use for what
uh you know for what task? Did you decide? Did I decide? No. The hugging face guys who built this particular
model decided that for us. Can you also choose which model to use? Absolutely yes. In this pipeline model, you can
change this model to any of the other models. And I'll also show you that like how can you pass your own model for a
specific task. You can definitely do that. You can change the model and you can perform your own classification as
well. But do you get that? This is the beauty of a model like this, a library like this, like transformers. You're
still working with code, but they've completely abstracted all of the open-source models that are available
right now. Text generation. This one seems familiar. Text generation. So I'm using this model and I'm asking it to
generate text. So the moment I say text generation, it is defaulting to the GPT2 model. The GPT2 model was an open-source
model back in the day. It was actually available for everybody to use. The GPT and the GPT2 models, both of these
models were free, were open source for everybody to use. only from GPT3 3.5 things started to
change because a lot of people tried GPT then came in things started to become very very complex that's when they
decided they're not going to open source the models anymore hugging face only stores it as a repository running um
that's exactly what I was showing you here so if you actually go up here if you go to models
hugging face has all of these models available with them on their cloud infra infrastructure. So if you want any of
these models, you can just go in here, you can download this model or you could just uh use some of this code and you
can also execute it. Remember I spoke about masked language modeling like of course I should also be able to do that
right just the way text generation I should also be able to fill a mask. This course will teach you all about mask
models in the AI space and I try to fill what the mask is. Um and I've asked it to make the top two predictions. So it
has used DL Roberta which is another model which is uh again hugging face has chosen this model for us. You can switch
to any other model you want. But it has predicted this as instead of mask it has predicted this as predictive models or
it also made a prediction of something called as role models which kind of is ridiculous.
All about role models in the AI space which um which sounds right but that's probably
not we intended. This course will teach you all about role models in the AI space. Well, role models in the AI space
is not a bad uh sentence. Grammatically, it's all right, but it's just that uh that's not what I am referring to, at
least in this context. Uh maybe predictive models is still not a bad uh yeah, just for it to predict the top
two. It has predicted top two here. This is what I wanted to show you all. I don't know if I had already shown this
to you all, but so what do you observe here? It's a very interesting visualization by the way. Uh so if you
see here 2017 was when the transformer models was published right it was around 2017 or or so the paper was published
and then from here what do you observe? You see 2019 BERT was published. Bird started to become very very popular
right encoder models right? So for generating any kind of embeddings using them for some specific uh you know tasks
all of that started to become very very popular here. So you see Bert, Distlbert, Roberta, exactly those are
the examples that we all using that I'm all showing you. By the way, those those tasks in the hugging face model, they
were all using these the same models. Suddenly it also used the GPT2 model. When the moment I spoke about text
generation, it went here. I said generate text uh text generation and it went to the GPT2 model. Right? It also
actually went for the BART model. It also picked up the BART model. When I wanted to to do like uh uh mask language
modeling, it went to the BART model as well. So my point is everything that you just saw is
basically this space. Everything in this area is what I just showed you. Everything here you have done in the
past. Word embeddings, wordtovec, glove, all of that. All of that you had already done, right? or the glove models, the
word toe models, the I think we we did you might have done this using genim or some other libraries in your sessions.
This I'm showing you how to use all of this through hugging face, right? How use some of these models through hugging
face. Then comes the next level which is this part. Then people suddenly realized that
you know what these transform models are doing very very well but they suddenly started to realize man if you actually
pump more data into it these models have a lot of capacity a lot of capability and that's when they started to create
larger models and thereby these large language models became more popular that's how the LLM then came into
existence The hugging face models or the models that we are trying to access through
hugging face, they were all open source at a point. The foundations of everything that we that that has become
so popular now. The foundations of all of that was open source was free for everybody to use, right? It was
basically active research. Companies like your u hugging face, the companies like your hugging face and stuff like
that. that is just uh you know that is providing all of these models for people to access. But suddenly what has
happened is we've suddenly progressed from here to here and the reason for that is because these transformer models
are so powerful that suddenly in 2021 22 later half of 22 and early parts of 23
people started to create more and more and more and more models. Right? If I actually create the same timeline for
2024, you would not believe how how long this this thing would be because there's so many models. That place is now so
cluttered. It's so hard for people to keep a track of. So the number of models that got published much later are so
many more and became very complex for people to start working with. Long story short that these language models that we
are trying to access uh you could have earlier also done them through hugging face. The thing is
hugging face has also still kept itself very relevant right hugging face has also still kept itself very relevant. So
what it has also done is even in this particular space all the models that are open- source right like the llama 2
llama models and so on and so forth it has provided all of that through the hugging face platform itself and it has
provided all of those models through the hugging face platform itself. So uh I'll quickly show you how you could maybe use
one of the models for a specific task like a classification or like a um a question answering. I'll just show you
how you could do it and then um that would already give you like a good idea of how you could use the same for some
of the other models. Um what is the difference between the left branch and the right branch? I we'll discuss that
ra'll discuss that not right away but we'll definitely discuss that. All right, let's go back to the hugging face
u tutorial. So by now we know how to perform all of these classifications. Text generation,
mask filling, question answering, all of that is available. Um
there are a bunch of encoder models, right? The kind of tasks that you can do with encoder model is things like these.
Here you go, Ragava. Encoder models can perform things like these, right? Sentence classification, named entity
recognition, question answering, extractive question answering, right? Extractive question answering can
be done using the encoder models, right? Your B and stuff like that. Your decoder models can do text generation. They're
only for text generation. You remember in the encoder decoder setup, in the encoder decoder setup, you
have on the left side you have the encoder, on the right side, you have the decoder. If you just use the left part,
if you pass a piece of text as an input, you generate numbers as an output, right? So you could use those embeddings
for doing any kind of classification, for predicting any kind of words, for predicting the next word, for question
answering and stuff like that. So you could do it for all of that. But if you just use the right part, right, if you
just use the right part, you could do it for text generation. Meaning if I just give you an initial word, if I just tell
you what the first word is, you can automatically start generating text. You can just take the second part and then
you can start generating one word after the other after that. Um that's basically text generation.
Encoder decoder you require both an encoder as well as a decoder if you want to do something more. Right? So if you
want to do things like summarization, if you want to do translation, if you want to do translation, of course you need a
encoder as well as a decoder. If you require if you want to do generative question answering which I'll explain in
a minute there also you will require an encoder and a decoder. So fundamentally that's how these three things are split
and that's why you see this tree as well taking a split like that. So encoder only, encoder decoder, decoder only.
That's how this is split. The some of the more nuance of how each one can be used for the others. I will explain in a
few moments because it has to you also will have to naturally grow into it. I'll explain that in a few minutes. See,
here's a here's an example, right, of loading an a a new model, right? So for example, if you want to load any of the
existing model, right? Let's take for example, let's take this this piece of this particular model, right? Dist base
uncased fine-tuned SST English. Let's actually go check take a look at this particular model on the hugging face
library. Okay, so this is the model that I just spoke about. So model description this model is a
fine-tuned checkpoint of dist bird based uncased fine-tuned on SSD2 which is a data set by the way this model reaches
an accuracy of 91.3. Uh what are the tasks you can use this particular model for?
Um you can do some kind of a classification using a model like this right you could
perform any kind of a classification using a model like this. Um so the question is how can you do that? So
here's the piece of code that they've provided for all of us to access, right? Um let me actually go back here. Let me
show you how you could do it. So I'm loading something called as automodel which is like a
wrapper, right? And I'm saying hey I want this model and then I'm saying model is equal to automodel dot from
pre-trained. So I'm saying hey fetch this particular model for me. fetch this as a pre-trade model for me. Some
weights of the model checkpoint are not used while initializing dist. This is expected if you're initializing dist
models the checkpoint of a model from another task. This is not expected if you're initializing dist from the
checkpoint of a model that you expect to be exactly identical doesn't matter. This is a warning more than anything
else. Uh but I've loaded this model. Now this model is available similarly just the way I load a model. Right? So
whenever you are dealing with um any kind of a um any kind of these models right what
you need to understand is you require two things you read a tokenizer and you require something called as a model.
What is a tokenizer? It basically takes a piece of raw text and converts that into input ids. So the
raw text goes into the tokenizer and the tokenizer token breaks the sentence down and it creates that into a bunch of
input ids, bunch of numbers, right? Uh it's like label encoding. Think of it as label encoding. That's
essentially what happens from here to here. Then these input ids are passed into the model. Internally embeddings
are generated and it generates predictions for you. It of course does not generate the prediction itself. It
generates logits which is basically the output from your before you pass it into a softmax or a sigmoid um you know so to
say right it generates the predictions you do some kind of basic pro postprocessing and it generates the
predictions for you. So you need two or you need three fundamental steps. Step number one is a tokenizer. Step number
two is a model. Step number three is pro postprocessing. Right? If you just look at tokenization itself, if you
fundamentally look at token and this is all by the way for any kind of a this is all applicable for accessing any model
um using um your hugging face transformers, right? You you access any model using it. It
the process is still the same. So you this is a good so this course is amazing, right? You pass that as a raw
text. So logits is this um so imagine you have a neural network okay uh and um the last layer let's say you are
performing some kind of a classification or whatever right so you have a lot of inputs and then finally you're doing a
classification in the last layer typically if you're doing a multiclass classification what is the u activation
function that you have on this what activation would you have if you're doing a multi multiclass classification.
If you're doing a multiclass classification, it is never a sigmoid. It is always a soft max. So how does if
you're doing a binary classification, you would apply a sigmoid. How does sigmoid work? You take sigmoid as an
example. 1x 1 + e ^ of minus whatever is the input that comes from here. I'm going to put h2.
Right? So the whatever input that you get from the previous layer, that's h2. Right? before applying the sigmoid
whatever output you get this H2 that's a logic right so the output before the sigmoid is referred to as a logic it's
the raw output then you pass it through some kind of a exactly it's a weighted sum from the previous layer your
activation from of the last layer is referred not the activation function but the activated output from the previous
layer is essentially referred to as a logit you put it through classifier um and you put it through some kind of
an activation function and get the probability itself just that part is referred to as a logit. Yeah. And and so
just double clicking a little bit into this tokenization, right? How does this tokenization exactly happen? So you take
this this this sentence, right? This course is amazing. The sentence is broken down. This course is amazing. And
then what they do is you remember when we spoke about encoder decoder I said you know
what we add like a start sequence and end sequence tokens. So you see this this is like a start sequence and this
is like an end sequence token cls and SCP. Uh I don't know what the CLS is supposed to even stand for. I think it's
a carriage line return that kind of a thing or whatever. I I don't even know what that is. SCP is a separation like
some kind of a separator. Exactly. It is like EOS and SO SOS and EOS start of the sentence and end of
sentence tokens. So two extra tokens are added. Of course, these tokens also have a default ID. So 102, this is one Z,
this is 101 and this is 102. And all the other numbers get the token ID, right? These IDs are static ids that are being
maintained at the back. Right? These ids, all of these ids that you're seeing, they are static ids. So these
static ids of course have their own uh associated embeddings. These ids have their own associated embeddings. That
embeddings will come later, right? So then once you get these ids, those embodings will be created against each
of these. All of that put together is then going to be passed into the model itself. Every ID is unique to that
particular word. Every word has a specific ID and every ID has its own embedding. So what are we doing here? Um
let's go back here. Let's see exactly how this works. So let's say we want to perform a simple by the way this
particular model is uh trademark the for the sentiment classification. I can pass it for any kind of a base.
The default uh use of this particular model is for a sentiment classification. Right? So I'm using this default model
for um uh for for for a for a for a simple straightforward sentiment classification kind of a setup. Okay.
But let's exactly understand how what happens inside. So here are two sentences that I'm passing. I've been
waiting for the hugging face course the whole life. I hate this so much. Right? There are two sentences.
I pass these raw inputs into the tokenizer. Right? The tokenizer I've also loaded. As I said I need two
things. I need the model. I need the tokenizer. So the I'm also loading the tokenizer here. But I'm using the auto
tokenizer class. To load the model I'm using the automodel class. to load the tokenizer I'm take using the auto
tokenizer class I'm just passing this as the key checkpoint and it is loading the respective tokenizer and the respective
model for me behind the scenes it has loaded it and it has kept it in memory for me right now now what I'm asking it
to do is I'm saying look take these sentences let's actually put it through the tokenizing and let's see what
happens what's what comes out as an output so if you observe it has taken this sentence it has broken it down it
has taken this sentence and it has also broken it down. Let's take the second sentence. Huh? This let's take this
sentence and it has also padded it for me, right? Uh and wherever the sentence size is large, it would also truncate
it, right? And it returns it as a PyTorch tensor. In this case, it returns it as a pytorch tensor. PT is pytorch.
So, wherever the length of the sentence is not enough, it'll pad it with zeros. Remember padding from computer vision.
Same concept here as well. it'll pad it with zero. Wherever the sentence length is uh larger uh it'll also truncate it
wherever required. So this sentence is a large sentence which is why it did not have to truncate. The first one is 101,
the last one is 102 here as well. First one 101, this is 102 and all of these are zeros. Basically indicating that
this is padding. And then I also have something called as an attention mask here. basically saying
on what should it actually compute the attention. Uh yeah, so there is a default length. The tokenizer has a
default length, right? So um if the length of the sentence is large, then it will truncate. If the length of the
sentence is short, then it will pad. If my sentence has if the expected sentence length is let's say uh 15, so whatever
the length of the sentence is here in this case. So 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 I think it is only 15 if I'm
not wrong. I don't know why it's 16. Um but if the expected length of the sentence is 15 tokens then it's 16. If
it is 16 tokens then after adding all the special characters if it is 16
then it will keep it. If it is less than 16, it will remove the additional characters. If sorry, if it is more than
16, it will remove the characters. If it is less than 16, then it will add zeros because I want both the sentences to
have the same length. No, I don't want the sentences to have different lengths as inputs.
Else I cannot add else I cannot club them as matrices. That's it. So I take these inputs, whatever inputs I get, I
pass them into the model. I get these out. I get these as the final inputs. So I get these input ids and I get the
mask. The mask is just to say on which word should I compute the attention. These are paddings. No, I cannot. It
makes no sense for me to compute any kind of attention on top of these words. I should just compute it on these words
and these. So which is why for the first one I have ones for everything. For the second one I just have ones on until
wherever the words are there. For everything else it is zero. Basically to say that don't use these words to
compute attention. computing attention on padding makes no sense that's why okay and what I simply do is I take
these inputs and I simply pass it into the output into the model the model generates the output for me it's saying
it generated the output for me in this particular case this is of course for the last hidden state um but it has um
generated the output for me in this particular case the vector output for the transformer model is usually large
it generally has three dimensions as you can see Here batch size as you can see here two b two units
sequence length the length of the numerical representation of the sequence 16 in our example in this particular
case it is 16 and hidden size the vector dimension of each model input. So the model has generated
the attention output for each for each input for each of these input words it has generated an output. So 768
dimensions is what it has generated. So 2x 16x 768 is what it has generated as an output. So this is highdimensional um
it's a high dimensional vector but it is a rich attention output is what it has. uh what we can then use is we can use
that for any of the other sub you know downstream classification models or any of the any of the steps that I want like
in this case right I can use that model as you can see here automodel for sequence classification I'm specifically
saying hey I want to do sequence classification here u and then I'm passing this model passing these inputs
and it is simply generating an output for me um and I can simply pre-process it post-process rather I'm applying a
simple softmax on top of it model passed the model into the I used the model generated the predictions you see these
predictions here I passed it into a softmax uh layer once I get the softmax layer first one is for you know point uh
004 and the second one is.99 so here's the thing I have the sentence right I have two sentences probability of class
0 probability of class one prob probability of class 0 probability of class one two sentences for each
sentence I have probability of class 0 1 0 1 so what I'm going to show you right now is uh I don't want to go here but
what I actually want to show you is the second part which is the output logic values can I explain okay great so take
a look at this so the model has been generated so remember so I have two input sentences so here's here's what's
happening Guys, so let me let me try and make it way super simple for all of you to understand. So I have an input
sentence, sentence one and sentence two. Two sentences. These sentences raw sentences. Step two, they've been
tokenized, right? And they've been when they pass when they go through a tokenizer, they actually become proper
outputs, right? So that that will generate an output of output matrix of 2x 16x
768. It'll generate a matrix of size 2x 16x 768. This is output from tokenization.
I'm sorry tokenization and embeddings. Actually let me just make a small change here.
This 2x6 is the output from tokenization. Then these 2x6 essentially meaning two
rows and 16 uh you know input ids each right is essentially passed as an input into the
next the model. Once you pass it into the model what does the model do? The model generates 768 embeddings or each
embedding for each of these units is essentially generated as an output. Right? The model has generated. So if
you take one sentence, right? Word one, word two, word one, word two, word three, all the way up
until word 76, sorry, seven, word 16. For each one, you now have a 768 long vector. Exactly. You have a
word embedding for each of these different words. That's essentially what you have. That's coming out as a model.
Now, what do you do? You take these embeddings, right? You take this particular
embedding, you put it through one more layer, right? You put it through one specific layer. Now, you typically call
that as a task specific layer. You take this particular thing. Now, what is this task specific layer? So, for example,
you want to perform and you want to perform let's say a binary class classification. So, what do you do? You
take this and then you add basically one more you add one more layer over here which then performs the binary class
classification. So when you do a binary class classification this 16x 768 will return what for each one you will return
two predictions as the output probability of class 0 probability of class one. By default it does not
generate a soft max or a probability right. So for that you need to add a last layer which is of a soft max which
will basically again give you a size of 2x2 but that will then finally give you probabilities. That's the difference
guys. That's step by step for you. These are the four sub subst steps over here. If you choose a when you go with this
auto model for sequence classification you're not loading a default auto model but you're loading that model for
sequence classification. So you don't need to apply any specific last layer computations here. You don't need to do
any kind of a you don't need to take the embeddings and again further pass it into another layer. The output itself
will be directly logits because it is already performing against a particular task itself for you. That's the only
difference here. But if you load the default model which is the default auto tokenizer just the default auto model
then in that case you will get the raw outputs and then you'll have to take these raw outputs and then perform
further uh one more layer of classification. This is the default model which will retrieve the hidden
states. But if you want to do for sequence classification for token classification for question answering
you have the same auto model for sequence classification auto model for token classification and so on and so
forth. So instead of if you simply use auto model for token classification, it will not generate uh embeddings for you
as an output, but it'll generate the the classification output itself as the final output. That's the only difference
my friends. Suppose instead of sentiment, if I want the prediction to say red, amber, green, can the
pre-trained bird model still work or will I have to be trained on specific data? You can use for zeroshot
classification. Suraj you could do zeros short classification here or you could also do few short the point is you could
do this take a sentence and predict it as as uh uh as this either you could do this or if you want to train you can
fine-tune the model you have a few examples you can fine-tune it right you can take this particular model and you
can fine-tune it you can say auto classification for sorry auto model for uh for for sequence classification you
can take this model and then you can fine-tune it. There is a fine-tuning class that is available and you can
fine-tune that particular model as well to for your specific data set. You could do that too. Let me show you one more
example guys. One last example but in this particular case uh I want to show you a question answering model. Actually
let's create it. No, let's what's the big deal about it? file, new file, Jupiter notebook,
PyTorch. Let me So, what we're going to be doing right now is
let's take one of these uh models. Let's actually go here. Let's go to models. Okay. My my task here is I want to
create a question answering model, right? Uh where is where is question answering computer vision sorry this is
here uh natural language processing question answering fantastic so I want to use any of these models let's
actually use the bird model itself let's use this model itself my objective is to be able to create a model which can take
a piece of context like this as an input which takes this particular text as an input or question as an input and
answers this question from this particular answers this particular question from this particular ular piece
of um context. That's my objective here. Okay. So the question is how can I create a model like this? Do I how do I
access uh you know or rather how do I create a simple model like this particular thing. Now for me to be able
to do that of course I would require um I would of course require to load this model right whatever that particular
model is. What I want to quickly also show you all here, let me just come back here. What I also want to quickly show
you all in this case is this particular model. By the way, this kind of an approach, there's a data set called
squad data set, right? The Stanford questionans answering data set. So, this is by the way the BERT model was
fine-tuned on this squad data set. What is the squad data set? Let's quickly take a look at it. The squad data set is
basically like a reading comprehension data set, right? So you all must be familiar with the reading comprehension
example, right? So reading comprehension is simply nothing but u you have reading comprehension is
simply nothing but you let's say have let's take any of this packet switching as a piece of text. So you have a
paragraph and then you have somebody asking a question and actually humans responding back to this particular
question. This ground truth is nothing but somebody there was crowdsourced. So somebody was somebody was actually
manually trying to respond back to these questions here. Um and then you also have something called as a prediction
which is essentially a prediction from this model the NLN net. So somebody actually made a prediction and that
prediction is what we are seeing. So you see these are all different models baff self attention
b single model. So if you go to packet switching these are the predictions from the b model and so on and so forth right
some of them it got right some of them it partially got right some of them it did not get right and so on and so
forth. So as you can see this is how the data set looks like. So the task is for you to use uh or whoever is building
this kind of a model, their task is to basically take any of the existing transformer models and then fine-tune it
on this so that you could do tasks like these where you pass a piece of context or a passage. You ask a question and you
ask ask it to answer that particular question given that particular context. That process is referred to as question
answering. That is the whole idea of question answering or that's know this is also shortly referred to as machine
comprehension. Instead of reading comprehension, you could also refer to this as machine comprehension.
That's how this whole thing works. Right? U now what we're going to be doing is we're going to be using one of
these fine-tuned models. Um and then we'll try and create an inferencing uh layer in our uh notebooks right at our
end. We'll try to quickly see if we can create a small inferencing piece of function using one of these models in
our notebooks. Right? So that is essentially what what we'll be doing. So that now what we will be able to do is
we can pass any context any piece of text and we'll be able to generate our own response from this particular
context. Um that is what we'll be setting up um in our notebooks for the in the in the next uh in the next few
minutes. Um so we'll be discussing the as I said u we'll sort of be doing a question answering example here um using
some of the stuff that we just saw. Uh right I think a lot of these things are pretty straightforward. uh we we we
could use uh you know some of these uh different uh topics that we just learned, right? We can try to do
question answering using the question answering pipeline or we could also do it using this uh auto model for uh for
for question answering sort of a setup as well. You could do either of either of these. Uh let's quickly do it using
the pipeline pipeline object right. So from sorry from transformers import
pipeline and we simply say pipeline is equal to pipeline of
uh the task is uh question answering I hope it's with the hyphen or with
underscoreen and then here I would also want to provide the model
um and the model here is in this particular case the uh what is the name of the model let's
go back to the model itself uh okay uh not this was the Google but right
let's go question answering we wanted to do this one. So let's copy the name. All right. And uh
once uh that gets loaded, which should basically be any moment what we could then do
is we could simply say ppl, which is nothing but the I'm sorry. Let's actually take a look at
what are the inputs that need to go into this. So the inputs that are essentially being
passed here are just trying to find the exact uh piece. Right? So the two things that should
ideally go as an input into both of these are uh let's go back here once again is the context.
It's the question and the context. There are two parameters that need to go in. Let me go back here again.
So, so there has to be a question and then there has to be a context. The context
is essentially a piece of text that I will be using to answer that particular question itself,
right? And the question is the actual question that I want to respond back. So, let's pick any uh Wikipedia article,
right? Let's go. Let's look at uh I don't know guys. Let's let's pick anyone.
Let's go. Let's pick any article. All right. From today's featured
article, let's read about it. I have no clue what this is. Uh Jersey Act was introduced to prevent uh registration of
most American bred thoroughbred horses in the British general stud book. There's a lot of uh lot of information
here. And um uh the loss of breeding records during the American Civil War. Actually, let's let's go here. Let's add
it. I have no clue what this is even supposed to mean, but I'm going to ask what is
Jersey? That's my simple question. And I'm going to say uh ppl
uh of question is equal to question context is equal to context and I'm going to say
result. Let's execute. Let's see what what happens with this. What I'm expecting is
this particular model try to you know tries to answer that particular question from the context that I shared here. And
uh let's look at the result itself. Yeah, here you go. The answer is to prevent the registration of most
American bred thoroughbred horses. That's not bad. That's a decent answer. I mean, I wouldn't consider that to be a
perfect answer, but it's a decent answer, I would say. Here's the part here's the point where I want to call
out one specific aspect, right? Um, what do I mean by it? So there are there are two kinds of
question answering that happen. So when you talk about question answering there are two kinds of question answering.
First one is called as extractive question answering and the second category is generative question
answering. What is the difference between the two? In the case of an extracted question, extractive question
answering, what you're simply doing is you're basically taking a piece of text and you're trying to answer that
particular piece of text from the existing context that is available only by extracting the relevant words. Right?
So what do we mean by it? In the case of extractive question answering, you essentially try to find a piece of the
answer in the existing sentence, right? It would say as you can see here, it has predicted the start and the end.
So it says the start is the 31st character and the end is the 100th character. So you go from left to right,
you find the 31st character, you start there and you go all the way up until the 100th character and then you simply
stop there. That's essentially what you mean by extractive question answering. So you're essentially extracting a part
of the existing context that you have provided. There is the other category of question answering which is called
generative question answering wherein wherein you are not trying to respond back from a particular question by or
respond back to a particular question by extracting the response but rather by looking at the complete question looking
at the complete piece of text and coming up with your own version of the answer. It may or may not exactly be present in
the in the context but you're essentially actually creating an answer in that particular case. That is
referred to as a generative question answering. Now what does that translate to in our world? What it should
translate to is what sort of goes into an extractive question answering and a generative question answering. In an
extractive question answering, you are just as I said just trying to find the start and the end of a particular set of
that response, right? Or other ways put you try to find the start and then you go you try to predict all the way up
until the end. That's how basically question answering works. You try to predict where the answer might start and
then go all the way till the end. So in a lot of way you're simply predicting the start and the end. That's all that
you're predicting for. You're predicting the start token and the end token. That's what you're predicting for in the
case of extractive question answering. However, in the case of generative question answering, you actually use the
complete encoder decoder architecture. You use the complete encoder decoder architecture. You pass the question, you
pass the context, both of them together as an input, right? Generate the embedding, pass the embedding, and then
you pass start of sequence as a token. And now you start emitting one word after the other, one word after the
other until you actually create the final response. So here in this particular case, you're
actually going through the complete encoder decoder setup. However, in the case of extractive question answering,
you don't need the decoder. You just need the encoder. You just need the encoder part of the transformer. Why?
You just pass a piece of text and the and the context as an input. And all that you're predicting here, you're not
predicting words. You're just predicting the start token and the end token. That's all that you're predicting for.
You're predicting for the start token and the end token. So in this case, just an encoder is enough. You don't need a
decoder. In the case of a generative question answering, you need the encoder and the decoder both,
which is why it is a lot more like the way you solve it is very different, which is why the kind of quality of
response is also very very different. This is referred to as whatever we just did is an extractive question answering
sort of an example. All right. Which is why also if you see here the encoder model here, if you see
the encoder model, it says extractive question answering. Encoder decoder together can do generative question
answering for you. Right? So back again, let's go back to this. Um so now at least you know given a particular piece
of information or given a particular piece of the um uh context you know how to do question answering but at least
the extractive question answering we know how to do um this of course we are using the pipeline
object. If you did not use the pipeline object, right? If you did not use this pipeline method, you know, method here,
let me show you how the code looks like, right? Let's let's actually look at how the code will look like if you did not
have that. So, had you not done this, okay, what you would have had to do is you would have had to do from
transformers or rather import automod sorry auto
model for question answering. And
from there you would have all of course had to also do from tokenizers
import automod for auto tokenizer for question answering.
Uh why is that the case? The models the package is not available. Auto
did I make a error somewhere? No, I just want to do it specifically for question answering. Ah, okay. Sorry, my this is
my bad. Okay, cool. Cool. Uh from tokenizers import auto tokenizer.
Hey, why is this not reading? Just a second. Yeah, that's another thing. Okay. And then you simply go the model
is the same. This is the checkpoint is the same. And then you say uh so the model is going to be automodel
for question answering dot from pre-trained. Uh and then you simply pass checkpoint
which is the model. uh and then you say tokenizer is equal to auto tokenizer
dot from pre-trained of ckpt which is basically both of them the models will be loaded right now and
once the models are loaded of course I can then take it one step further I can use both the question and the context
I'll use the same question and the context here and Now simply going to say hey look the inputs are going to be
tokenizer of remember this is the tokenizer so questiona
text I need to pass both of these return tensors is going to be let's see if it can return an numpy sorry this is
context this is not text let me stop this and Uh let's go. Let's execute this.
This will take a couple of seconds for it to get executed. All right. Perfect. So now this has been executed. Let's
quickly go to the inputs. Let's see. So my inputs are now been created. So if you just look at uh this
is by the way the numpy arrays of actually let me just go back to PyTorch.
So I think it returned it as numpy arrays. So let me just return it as uh pytorch so that I can then pass it into
the model itself. So now if you look at the inputs the inputs have now been returned
with the this is nothing but so if you see here if you look at the tensor
it has you see the 101 all the way up until 102 but I'll just
quickly highlight one part wherein you see a quick difference. See look at this. So 101 all the way
till 102 and then you have the rest of the same thing all the way up until 102. Why is that the case? So the 101 all the
way up until 102. This is the question that's the question and from there this is the response or rather this is the
context. So both of them are sort of combined together. So if you see which is why if you see something called as
token type ids right if you see the token type ids there are a few zeros and then the rest are ones the zeros are
nothing but your questions the ones are the context so it is a way to tell your model that look when you generate the
response only generate the response against the ones don't generate the response against the zeros these ids are
predefined I can actually show that to you as well these are the pre these are the ids that
are that the model actually bears with it. Let me actually show that also to you. You can say tokenizer dot convert
ids to tokens. So what you could do is you can take these ids. So this particular tokenizer my friends is a
very very specific tokenizer for this model alone. Remember that this bert model is a large
model. It has been trained on massive corpus of data. Right? It has been trained on lot of data in the past. Uh
Google news this and that. It's been trained on a lot of the corpus. So what you could possibly do is you could
simply as you know take these ids sorry input input inputs of
input ids. So you see these ids, I can actually pass these input ids as tokens here and I can ask it to generate the
whole thing. Uh okay pass this.
See there you go. The moment I pass these tokens into
I pass these exact tokens back into this function called convert ids to tokens. Look what it has done. It has actually
given me how the sentence has been broken. What is Jersey act question? So each of this is a is a token separator.
The Jersey Act was introduced to prevent the registration of most American hyphen bread thoroughbred horses in the blah
blah blah and so on and so forth. So this is how it has actually been tokenized and these words my friends are
exactly the same words that would have been also tagged against these ids when the model
was first trained on its large data set. The model would have been trained on a very large data set, right? So it would
have been trained against a particular set of ids. Every word would have been mapped against a particular ID. All of
those are going to be here. In certain cases, what may happen? There might be words here that might not have that
might not have been seen otherwise. Right? In such cases, they will all be tagged as something completely random.
So for example, if I say if I have a sentence like that now let's see what happens
either this should get broken down into individual tokens that if the tokenizer is smart or it would simply tag it to
something random which I would guess it should now see what it has done. It has taken
that particular sentence or that particular token and it has actually broken that particular token down into
smaller tokens because the tokenizer doesn't recognize this particular word. What it has done
is it said okay I don't recognize this word. The tokenizer has this functionality of also trying to break
that particular token down or that particular word down into the nearest tokens that it might have seen in the
past that it may have seen in the past. So so that it can at least match for some
similarity there in that case instead of simply saying I don't know what this word is. Does that make sense to you?
Is this clear everyone? Cool. Fantastic. Now, now that you know how the tokenizer
works here, now that you know how the tokenizer has created all the tokenizer and all of the ids, what do you do then?
You take the inputs and then what do you do? You pass it into the model, right? So, you say, hey, no, first you
need to generate the output, right? So output is equal to model of star star inputs
how many other inputs you have. In this case it's only one input but I'm just passing it anyways. Um so now if you
look at the outputs so look what the output has uh created. So question answering model output start
logits. So now if you observe it has given me logits for the start and the end logits. So if it has given
me start token logits and it has given me end token logits basically I need to then find of these
the one that has the highest logit and then work from there. Does that make sense to you everyone? Do you
understand? So from here I need to find okay the logit that has the highest value highest logit that would become my
start and the logit that becomes my end that would become my end. then I need to use both of these and stitch the answer
together. Let me show this to you quickly. Let me quickly show you. By the way,
I'll just remove this one. I I don't want to confuse the model. Okay. So now let me generate the output.
So I have the outputs. So let me show you what what we'll do from here. So now what I'm going to do is I'm going to say
output of I think start logs. So you see the start
logs it's a tensor right and from here I'm going to say um let me just import numpy or pytorch
also should work but Okay, this is the start logs. Okay. Uh
let me just say with no grad no grad. Sorry guys.
The torch ngrad is essentially a way to say that hey I don't need this these tensors to go through any more uh
gradient computations especially when you're doing inferencing you don't need to do that uh now it'll
just convert get converted into a numpy array. So from here all or I can also do a simple argmax.
So it says hey it's the 12th tensor it's the 12th token that has the highest value. Uh in this case that's the one.
Similarly the end logits. It's the 28th one. So the start is the 12th and the end is the 28th.
And all that I need to do now is I sort of will have to simply all that I have to do is I simply will have to you know
run through the uh you know run through the input ids and then go one after the other.
Right? So I'll have to simply say uh where did this go? This is the start index.
This is the end index. Right? And all that I need to do now I simply will have to say um
whatever inputs I have which is nothing but this of
input ids. These are my input ids. Right? All that I need to get I need to get the start
index. all the way up until the end index. Right? I need to go from the
start index all the way up until the end index plus one so that I also
cover that part. Um what is the start index and the end index?
Okay, there you go. That's it. That's the response.
I need to get the zero because uh without that it it was indexing the other one. So that's it. These are the
tokens, right? So these are the output tokens. These are the response or rather
output ids are these. And what do I need to do? I need to put these output ids through
this function. Remember this one. All that I need to do is I just need to do that.
That's it. There you go. That's the response. So I can simply say
that's the response. You would have had to do all of this or you could just use this or you could just do this.
Right? So this is method one. Whatever you saw here, right? Or
this is method two using
this and Yes,
you could use anything that you are comfortable with, anything that you would want to use. That's totally up to
you, your choice. Of course, you could either use the auto model for question answering and auto tokenizer, generate
the predictions. This would give you a little more clarity exactly what you're doing. you'll have to do all of this
extra stuff um which is sort of you know indexing extracting and and all of that. Uh but but this is also actually not a
bad idea because it might also be a good uh exercise for you all um you know to exactly see what goes first, what goes
next and so on. Uh or you could just use the pipeline object and pipeline will take care of all of that under the
hoods. That's the beauty of the pipeline object. the pipeline object will take care of nicely putting all of this for
you together. So it's this is so much more simpler that way. Cool. So congratulations. So that's how you go
about doing any kind of question answering using the BERT model or I mean question answering is just one task. You
could use you could basically follow the same approach for any other task that you wish to. Okay. Now now that we
understand all of this, let's go one step further. So, so this is great, right? I mean,
everything was nice and rosy all Yeah, don't worry guys. I I'll provide all the code that I have.
So, so do not worry about that at all. I'll definitely share all of it now. So, what now? Right. So, this is great. Um,
overflowing from your mind. Yeah, sure. Uh, cool. I I'll share this as well anyways.
So, so don't worry about it. Now,
this is all great. Now, what? Right. So, let me just open this. So, we understand how transformer models
work. We've tried out these transformer models through uh you know through through the hugging
face uh you know interface. Now the next part which is this is all great what suddenly happened right so so
this was nice there was still some skill in the game for data scientists here until this particular point but now what
happened after this completely threw the data scientist under the bus right so it simply said look man we don't need you
anymore um because these models have now suddenly become very very they have suddenly become extremely
powerful. So what happened there? What happened was that large language models suddenly came into existence. So people
realized that these transformer models had a lot more to offer, right? So they they simply just did not stop there.
They of course said okay let's use the the same encoder decoder architecture, right? So the same encoder decoder
architecture. Let's take the same models. So this so that so the same transformer models with attention.
Let's actually pump in more and more and more data to it. Right? So they started pumping in more
and more and to their surprise these models started getting better and better and better. They just started getting so
good that more data simply meant to better models. Right? That is where that took
us to this very interesting uh space called large language models. Right? So then instead of just 330
million parameters, you know, 250 million parameter models, you suddenly start started to see 7
billion parameter models, right? You have you suddenly started to 7 billion, 8 billion, right?
uh the number of parameters just suddenly shot up and that's where things started to
become super interesting. Why? Because these models have now pretty much become like they've become
so good that they understand languages like never before, right? Uh even better than let's say these BERT models so to
say. The B models were good. There was nothing wrong with it. The BERT models were very good. the GV2 models were fine
but but the beauty of the transformer architectures just made them made these models so much but so much
better uh moving into the the moving into the next area. So now what all of these folks started doing is they you
know be it be the be it the uh the opening eyes of the world or the Googles of the world or the metas of the
world they started training larger and larger and larger and larger models that's where
you we started to see this whole new branch come through called as these large language models. Now to be very
honest, large language models are just another another part of transformer architectures. It's just another
transformer model. But what but but the but the way they've been trained these large language models suddenly starting
from the you know starting from the the launch of GPT3
uh you know or or rather to be more precise Chad GPT so to say since the launch of Chad GPT
the world was taken by a storm right so this was in November 2022 was when November December 2022 was when
this particular thing happened uh chat GPT was launched within no time people started to use this this this this
capability uh and it start it saw like this insane adoption in terms of uh in terms of
usage everybody started using it the model started to become better and better and and the way they open sort of
set it up was also very interesting they kind of set it up in a way that more the more the people started using it the
models actually started to get better and better. So what is it that these guys have done
right let's let's take specifically the case of let's say a chat GPT what have these guys done which is significantly
different from that of let's say any other transformer model is it the model itself that has changed or has there
anything been different or have these guys approached this whole setup very differently the answer is actually the
latter right the model has not significantly changed. Of course, they've trained it with large volumes of
data. We still to be very honest don't know how ch you know the models powering chart
GPTs were trained. We still have no clue, right? We still have no idea how a GPT3.5 was trained. We still have no
clue how let's say a a GPT4 is trained. But we know broadly how the setup sort of looked like, right? Uh let me
actually bring up the uh one of the slides that was presented by the person who had trained the GPT3.5
models. Let me let me actually bring that up. Uh okay. So let me start with how
uh you know some of these models are trained right the size of of these models right size of the data that sort
of goes into this this is like a good example of u how one of these models was trained let me just uh this is actually
not the GPT model this is actually the llama model but still you get an idea of broadly how how they're trained so this
is common crawl. Common crawl for all you know is like a co Yeah, common crawl is basically like a very popular uh is
like a publicly available scrape of the internet, right? It's like a prawling engine that was built by uh
you know some very very novel normal you know uh very very noble folks out there uh who wanted to make uh data available
for everybody to access. So this is by the way common crawl um it's an open repository of all the web crawl data
that can be used by anyone. So you and I, anyone can download this data. You can use it for whatever you want. It's
for your consumption, right? This is a data set that is available since 2007. And uh it's crazy, right? The amount of
data that's available, it's it's sort of crazy over here. So if you look at uh statistics just to want just for you to
understand how big this data is um let's actually go. Yeah. So if you look at the size here uh I wanted to show you one
particular part which is the size actually yeah I just want to show you the size but anyway actually you can we
can actually see it here as well that doesn't matter let's actually come back here so if you look at this so almost
3.3 terabytes of data is what we're talking about that's how big this data set is for training and learning isn't
uh data sets in Kaggle enough no man I mean this is just to let you know right so the common crawl databases 3.3
terabytes. This is the whole of the internet, publicly available internet. It's as big as that, right? And of
course, there is no way you and I can use any of that data on your on your machine. So 3.3 terabytes,
right? And then there is bunch of others. GitHub code 328GB, right? All of the publicly available
GitHub. Wikipedia 83GB. Wikipedia is as small as that. stack exchange, right? So your your archive which is nothing but
your publicly available papers 92 GB books 85GB worth of books and so on and so forth. So if you actually take all of
this data, this is massive. So a good 67% of the model has been trained on almost all of the internet, right? And
if you look at the number of epochs, which means that the model has barely seen all of this data once, GitHub
barely seen 6. Only 64% of this complete data has actually been used during training, right? So Wikipedia, books,
and so on and so forth. the model actually hasn't even seen all of this data or maybe has barely seen some of
this data at least once. So the point is that's how big of a data are these guys using to train these models. So imagine
it's it's it's way way beyond your and my uh you know uh from an accessibility standpoint is just you
have the access to this data but it's just impossible. I mean the amount of resources that is required to train
something like this is crazy. That's how much data you really need uh to be able to train these models. And the beauty is
the large language models or the transformer architectures are actually able to consume so much amount of data
and give out something extremely good. So that's the more interesting part here. It's not to say that the model is
uh you're just pumping in so much data. The model is actually able to also consume so much data and extract some
very very valuable insight out of it. And then people have started to make some very interesting changes to this
particular model. What have they done? So what they've done is here the first to start with right u this pre-training
is like your regular model pre-training right so language modeling predicting the next token. So the raw internet was
taken. Basically everything that you see on the left side. All of this information was taken. That information
was passed into the model to predict the next word or the next token. That way you were able to build the base model
using a regular transformer architecture. Of course this is the GPT you know setup and these guys may have
used something very very specific that you and I probably don't know. they might have come up with some
innovations, some inventions under the hoods that you and I probably don't know and that's okay. Uh broadly it is still
the transformer architecture that much we know. Then what they've done is to be able to train this model itself they
required like thousands of GPUs, months of training and this is basically all of your GPT models, your llama models, palm
models, they are all basically the same category of of models so to say, right? That's essentially the pre-training part
of it. Then comes the second part which is the supervised fine-tuning. Now when we access the BERT model my
friends right that is essentially this part that is essentially just this part right we just had access to this
particular model nothing we did not have to do anything beyond this what now people started to do is they started to
not just stop here right but they started to take it one step further what they started to do is they started to
now train this particular model on something called as an inst instruction set right now what they've started to
ideal assistant responses so now they started to say hey look given a particular question
what is the ideal response that I'm looking at remember in the first go on the left
side the model was simply pre-trained to predict the next next token that's it was not doing anything else but here in
the super fine fine supervised finetuning to assess how good the model because you always have the raw data.
You have this sentence. If I am predict trying to predict given all of this, if I'm trying to predict this next next
word, I can predict and I can always compare, right? Because all of that data is already available. So it's not that
complex for me to predict the next token and compare. But the complexity sort of kicks in here, right? Where you need
ideal assistant responses. So what they started to do is they started to create questions or rather prompts right and
associated responses to these prompts right and they're saying hey given this particular prompt this is the best
quality response that I'm looking at this is the ideal response that I'm looking at like an assistant and as you
can imagine this information has to be manually written right so nobody might have this uh somebody has to actually
curate these responses so as which As you can see, written by contractors, low quantity, high quality. So, they're
actually fewer in number, but they were written by specialists. They're fewer in number, but they're very, very high in
quality. This is to tell the model how it needs to respond back as an assistant. So, that is what happened
here. So, they used that and they then further used that to further train this particular model. So, this part is done.
So you trained the raw model using all of the internet. Here you have trained it to do something very very specific.
So you fine-tuned that model to do something very specific. So until here it is the regular BERT like model. It is
until here it is regular GPT like model nothing change nothing different here. But here is where things started to get
better. Right? Now from here what they started to do is they started to do something called as reward modeling.
Right? Again 100,000 1 million comparisons written by contractors low quantity high quality.
What they now started to do is they started to evaluate all the responses that are given by the model if they're
good or bad. So the thumbs up, thumbs down, you might have seen on the on the on the chat GPT thing, the thumbs up,
thumbs down before they even actually put it out there, they started to do it themselves internally with a lot lot of
contractors. So they can predict given any particular respon or they can actually predict given a particular
response, is this actually good or bad, right? Is the response actually good or bad? Given
the question they're trying to predict, is it actually good or bad? Good or bad? Good or bad? and so on and so forth,
right? If it is good, they went ahead. If it is not good, of course, they they further went back and they fine-tuned.
Then comes the last part wherein the reinforcement learning sort of then kicked in. What is the last part of
reinforcement learning? This is where again now that I know that the response is good or bad,
can I actually go back and adjust my response to make the response better? So that is referred to as reinforcement
learning. RL HF reinforcement learning based on human feedback. Right? So I it learned a reward function it learned
that ah you know what if I do this then it is good. If I do this it is bad. Now the model knows that given a particular
question and a response as a user would the user like it or not. Right? Now the mo I've trained a specific model just to
