Why Probability? Handling Uncertainty in AI
The lecture begins by addressing a fundamental limitation of classical logic and deterministic AI: the real world is rife with uncertainty. We need probability to model:
- Partial Observability: We rarely know the complete state of the world.
- Noisy Sensors: Our perceptions and measurements are imperfect.
- Model Limitations: We cannot process all data or know every rule of how the world evolves.
Probabilistic assertions effectively summarize our ignorance and laziness, allowing us to build agents that can make rational decisions under uncertainty.
Core Probability Concepts
1. Possible Worlds and Events
- Possible Worlds (Ω): The set of all atomic, mutually exclusive outcomes (e.g., a die roll: {1, 2, 3, 4, 5, 6}).
- Probability Model: Assigns a non-negative probability to each world, with all probabilities summing to 1.
- Events: A subset of possible worlds (e.g., "roll is odd" = {1, 3, 5}). The probability of an event is the sum of the probabilities of its constituent worlds.
For a deeper dive into foundational ideas, see the Introduction to Probability and Statistics: Key Concepts and Terminology.
2. Random Variables
A random variable (RV) is a function that maps each possible world to a value. It's a deterministic function of a random outcome. Key points:
- Domain: The set of possible worlds (Ω).
- Range: The set of values the RV can take (e.g., {true, false}, {hot, cold}).
- Probability Distribution: Gives the probability for each value in the RV's range, defined by summing probabilities of outcomes that map to that value.
3. Joint and Marginal Distributions
- Joint Distribution: The probability distribution over multiple random variables simultaneously (e.g., P(Temperature, Weather)). It contains all information about the relationships between variables.
- Marginal Distribution: The probability distribution of a single variable, obtained by summing out (marginalizing) the other variables from the joint distribution.
Expert Insight: You can derive marginals from the joint, but you cannot derive the joint from the marginals. This is a critical point that underlies many challenges in probabilistic modeling.
4. Conditional Probability
Conditional probability updates beliefs based on new evidence. The fundamental definition is:
P(A | B) = P(A ∧ B) / P(B)
- This is read as "the probability of A given B."
- It renormalizes the joint probability to focus only on worlds where B is true.
- A conditional distribution (e.g., P(Weather | Temperature)) is a separate probability distribution for each value of the conditioned variable, achieved by normalizing the relevant column/row of the joint distribution.
The Chain Rule (Product Rule): An essential identity derived from the definition of conditional probability:
P(A, B) = P(A | B) * P(B)
This can be extended to an arbitrary number of variables, allowing any joint distribution to be expressed as a product of conditional distributions.
For a visual example of these concepts in action, see Calculating Conditional Probabilities Using Tree Diagrams and Dice Rolls.
Making Inferences: Bayesian Reasoning
Inference by Enumeration
A brute-force method for computing a posterior distribution:
- Select entries from the joint distribution that are consistent with the evidence.
- Sum out (marginalize) all hidden variables.
- Normalize to ensure the resulting probabilities sum to 1.
Problem: This approach is intractable for real-world problems due to exponential growth in the number of possible worlds and the amount of data needed to fill the table.
Bayes' Rule: The Foundation of Modern AI
Bayes' rule is a fundamental tool that allows us to compute the probability of a cause given an effect, even when we have data only about the effect given the cause.
P(A | B) = [P(B | A) * P(A)] / P(B)
Practical Example: Meningitis Diagnosis
This example demonstrates the danger of ignoring base rates:
| Variable | Probability | | :--- | :--- | | Prior probability of Meningitis (M) | 1/10,000 | | Probability of stiff neck (S) in general | 1/100 | | Probability of stiff neck given Meningitis | 80% |
Using Bayes' rule to find P(M | S):
P(M | S) = [P(S | M) * P(M)] / P(S) = (0.8 * 0.0001) / 0.01 = 0.008
The probability of having meningitis given a stiff neck is only 0.8%, despite the strong link between the two. This highlights the crucial role of the prior probability.
The Power of Independence
Independence
Two variables are independent if learning about one never changes your belief about the other.
- Formal Definition: P(X, Y) = P(X) * P(Y) for all values of X and Y.
- Implication: P(X | Y) = P(X)
- Benefit: Allows for an exponential reduction in the size of the state space. For n independent coin flips, you only need n probabilities, not 2n.
Conditional Independence (A Teaser)
True independence is rare. The most powerful and common concept is conditional independence, where two variables become independent once a third variable is known. This is the core principle behind efficient probabilistic models like Bayes' nets.
Example: Traffic (T), Rain (R), and a person carrying an Umbrella (U).
- T and U are not independent (knowing one can inform you about the other via R).
- However, given R (P(T, U | R) = P(T | R) * P(U | R)), T and U are conditionally independent. Knowing about traffic gives you no extra information about the umbrella if you already know whether it's raining.
This principle forms the basis for the next lecture on Bayes' nets, which provide a tractable way to represent and reason about complex probabilistic relationships. For a broader view of how these ideas fit into AI, explore Probability & Statistics: The Ultimate Guide to Modeling Uncertainty (Course Intro) and Comprehensive Introduction to AI: History, Models, and Optimization Techniques.
Okay. I'm going to get started on a new topic. Probability.
I hope that for a large majority of you, a large majority of this is quite familiar to you. But we're just going to go through the foundations
of probability anyway just to make sure that it's fresh in your mind for when we get to the more complicated stuff.
So, this is an introduction to probability. And then from there, we're going to go into Bayes' nets. Cam is going to be covering that starting next week.
And so this is just the main tools that we'll be using there. And the reason that we can't just do everything with logic is that the real world is rife with uncertainty.
For example, if I leave SFO 60 minutes before my flight, will I be there in time? Well, maybe.
You can't just deduce that from things about the world. You can't deduce it one way or the other. Sometimes you will, sometimes you won't.
So here are some problems with our understanding of the world that lead us to need some way of quantifying our uncertainty. So, we don't know the whole state of the world
in everything, I think. Yeah. In pretty much everything we've talked about up until now,
we know the whole state. When we were doing search, we know we're in this state, we know we're trying to get to that state,
we know the states along the way, we know how our actions take us to different states. In the real world, we have no idea what the state of the world
is. The state of the world contains every single fact about the world.
I don't know how many people are on the 7 bus in the capital of Nigeria right now, assuming such a bus exists. There are-- the number of things I
don't know about the current state of the world is much more than I could begin to describe. So we have to deal with partial observability
where we observe some pieces of the world and we are uncertain about the rest. We have noisy sensors.
I don't have perfect vision. I can-- let's see. I mean, I know that that red thing out the door
probably says Pull Fire Alarm or whatever, but I can only just see little bits of white on it. And so we have to manage our uncertainty
about even the things that we see or hear or whatever. Or more distant sensors like my sense of the traffic on the Bay Bridge given a radio traffic report
is a little noisy. I mean, they'll be telling me about the traffic at one point in time, I'm interested in the traffic
at another point in time, so you could say that my sense of what the traffic will be at that later point in time is a little noisy,
it's not perfectly accurate. And so those are fundamental problems. And then there are practical problems.
You could also call them fundamental if you like. There are practical problems, which is that you can't possibly--
I mean, even if you got all the data, there would be so much data about everything in the world. The state of the world, even if it
could-- even if knowledge of it could be brought to me, I could never begin to process it all. And lastly, I don't completely how the world evolves.
I don't know exactly how a tire driving over a nail will cause it to puncture or won't. There are so many details that are
relevant to the probability-- or relevant to the possibility that I get to my flight in time or not and I just couldn't begin to model them all.
So we use uncertainty not only to model real uncertainty in the world to the extent that exists, but also, and maybe mainly, to model our own limited ability
to comprehend it. So probabilistic assertions summarize all the effects of ignorance and laziness.
Since I don't know everything, I just say, okay, well, maybe this, maybe that, I'll just say probably this, maybe that,
and that's papering over a whole lot of things I don't know. And what we'll do after this segment on Bayes' nets is then we'll get back into decision-making
we did with search at the very beginning. and we'll tie it together with utility theory and decision theory and use it to make agents that maximize expected utility.
That is, they find the action that optimizes the quantity of the utility weighted by-- the utility of a given state weighted by the probability
that we get to that state. But that's looking ahead, so let's get to the basics first. So, what is probability?
We begin with a set big omega of possible worlds. For example, if we have a die, the possible worlds are it lands 1, it lands 2, so on to 6.
Six possible rolls of a die. There they are. That's our set of possible worlds.
A probability model assigns a probability-- a number-- to each world little omega, each possible world. So each one of these is--
did that-- didn't have any effect out there. Each one of these is a possible world that little omega could refer to.
So for example, in the dice world, maybe all of these outcomes have the same probability of one-sixth.
Could be a different probability model that assigns them different probabilities, but if we have a fair die, that would be sensible probabilities
to assign them. And these numbers, these probabilities that we assign to different outcomes,
have to satisfy two things. They have to be non-negative, first of all, and they have to sum to 1.
So also, none of them can be greater than 1. So each of these is an outcome. Those are the atomic possibilities.
You can't divide them any further. And then event-- an event is a set of outcomes. So it's a subset of the set of all outcomes.
So it's a set of outcomes. So for example, the event that the roll is less than 4 would correspond to the set of outcomes
or the set of possible worlds 1, 2, 3. Or the roll is odd corresponds to the set 1, 3, 5. So here they are depicted.
These are the elements of the event less than 4, these are the elements of the event odd. And if I'm not--
so just raise your hand high because-- yeah, feel free to interject with a question, but I might not see it.
Okay. So, the probability of an event-- that is, a set of outcomes-- is going
to be the sum of the probabilities over those outcomes. Outcomes, possible worlds-- here, it's shortened to worlds,
all mean the same thing. So the probability of the role being less than 4 is the probability of each of these elements
summed together, which sum to 1/2 in this example. And here's a cool, spicy little fact for those of you who think that this is just one way of doing things.
No, it is the way of doing things. If you do these-- if you do this in any other way, you can be exploited.
So if you make bets according to probabilities that don't obey the laws that I've told you, then someone can offer you a sequence
of bets that will guarantee that you lose money to them. Now in reality, you might find this suspicious and stop making bets, but the fact
that you should be making bets according to your own idea of how probabilities work, and yet they would lead you to a certain loss should
be a bit worrying to your idea that you're doing things in a reasonable way. Unfortunately, we don't have-- yeah, but questions?
STUDENT: What did you mean by [INDISCERNIBLE]?? PROFESSOR: So let's say that for an event was a set of outcomes that included A and B, and the probability you assigned
to the whole event did not equal the sum of the probabilities you assigned to each of the two outcomes within the event. So that would be--
if you assign probabilities in that way, that would be in violation of this rule. It would be a weird thing to do.
So maybe you're skipping over the possibility in your mind because it would be so weird, but it is a possibility, and if you did assign probabilities
that violated those laws, then you could-- then that could lead to bad consequences. Unfortunately, we don't have the time
to prove this in this class. Gets a lot-- gets into a lot of math, but it's interesting. Okay.
The next basic concept, when we're talking about probabilities, are random variables. A random variable-- we'll be talking-- we'll
be using a lot of random variables coming up. They're usually denoted by a capital letter. And it's some aspect of the world about which we
may be uncertain. So we could think about the random variable which corresponds to the number of people that
show up in any given lecture. Every time it's a little different. Nobody can really in advance what it's going to be.
And technically, it's actually not random, nor is it a variable. Sorry about that.
Formally, it is a deterministic function of an outcome. So we have our set of outcomes, big omega. We have individual outcomes, which we'll sometimes
refer to with a little omega. And our random variable-- what our random variable does is it takes in any of these outcomes
and it spits out the value that it's going to take if that's the outcome. So, for example, let's say we have the random variable odd.
And this is a function from our outcome space. So 1, 2, 3, 4, 5, 6. Two of the values, true or false.
So this is a Boolean random variable. So when we apply this function to the outcome 1, it returns true, and we apply it to the outcome 6,
it returns false. So you can see how this is a deterministic function, but you can also think about it as something
that varies randomly because depending on what the dice roll is, the input is going to be random, and so the output is going to inherit that randomness.
And so if we're writing the probability that it's odd, we might write the event the probability that odd equals true, the probability that the random variable takes
this value. Or we might just shorten that to the probability of odd using some of our logic notation also.
Another example, let's say the temperature can take two values. So this is a random variable. The domain is always the outcome space, big omega.
The range in this case is this set of two elements, hot and cold. So the random variable can take either of these values,
and we can ask, what's the probability that T equals hot? What's the probability that T equals cold? The range of a random variable can also
be-- it doesn't have to be finite. So the random variable that says how long will it take to get to the airport, that could take a value anywhere
from 0 to infinity. Or location of the ghost, that could be a random variable with a range that is a set of
ordered pairs of integers given-- yeah, indicating the possible locations it could be. So random variables will have probability distributions
associated to them, and these are defined as follows. So the probability distribution of a random variable gives the probability for each value in its range.
So the probability that the random variable takes this value or that value or that value. Basically, the probability, when you put a random outcome in,
that the output of this deterministic function is this or that or the other thing. So, the probability that big X equals little x--
this is our random variable, this is a particular outcome-- is the sum of the probabilities of the outcomes for which the random variable outputs that little x.
So here's our function, our random variable. It takes in a possible world, it takes in an outcome. Sorry.
And sometimes it'll output little x. Was that a question there? No?
Okay. Yeah? STUDENT: [INDISCERNIBLE]
PROFESSOR: Opposite. That's the random part. So this function is never changing, really.
But the input will be randomly generated, basically, according to this probability distribution. So unless you were saying the set itself is fixed.
The set itself is fixed, but when we select an element from a distribution, that's where the randomness comes in.
So this is basically like sample an outcome from our outcome space according to their probabilities. And this-- and then we check to see which of these outcomes
map to different possible values that the random variable could take, and that's the probability of our random variable. So if we wanted to know the probability of odd,
then we look at all of the outcomes such that the random variable outputs-- such that the odd random variable outputs true.
That's 1, 3, and 5. Each of those have a probability of 1/6. We add them together, we get 1/2.
And you can abbreviate this expression P of X-- P of big X equals little x as just P of little x. And then this refers to the whole distribution.
So this would refer to-- it's like a table of all the outcomes and all the probabilities associated with them.
It's a bit of notation. So just-- you might have to refer back to this if you forget what the notation means.
So, let's look at an example. We can think of the temperature as a random variable, and it could output either hot or cold.
So here's the temperature and here's its probability. Maybe it's 50/50 hot or cold. We can look at another random variable-- maybe the weather.
Capital W. Here is the probability that capital W equals sun. Here's the probability capital W equals
rain, fog, meteor listed over-- listed to the right over here. So sometimes we exist in worlds where multiple random things are happening at once.
There is a temperature outside. There is weather outside. And we would like to know not just individual probabilities
about them in isolation, but how they relate to each other. So this is the joint distribution of temperature and weather.
And then we assign a probability to every possible outcome here-- assign a probability to the possibility that it is sunny and hot, sunny and cold, rainy and hot,
rainy and cold, foggy and hot, foggy and cold, meteor and hot, meteor and cold. And when we have a joint distribution
of multiple random variables, the random variables that compose it are called the marginal distribution. So here's the joint distribution,
here are the marginal distributions. And some important thing-- one important thing I want to note here because maybe it
looked like I could just derive that from these two things, that's not the case. You cannot deduce what the joint distribution is from
the marginals, but you can deduce the marginals from the joint, and I'll show you about how to do that in a second.
I think it's very intuitive. So now let's talk about making possible constructing a formalism for a possible world given the things
that we're going to want to measure about it. So often what we can do-- and what we'll do in most machine learning contexts, is we'll begin with the random variables.
And this should say ranges, not domains. We'll begin with the random variables and their ranges. And we'll construct possible worlds as assignments of values
to all the variables. So we knew we wanted to think about temperature and weather, so then we just constructed this set big omega of outcomes,
eight possible outcomes, one for each combination of weather and temperature. So we'll just a possible world for us
will be an assignment of values to all of the random variables. So for example, if we have two dice rolls, roll 1 and roll 2, how many possible worlds are there?
Anyone can raise your hand. Yep? 36?
Yep, six possibilities for the first dice-- for the first die, six for the second, and any pair of them will work. And who can tell me the probability
of those possible worlds or outcomes? Yeah? 1 over
STUDENT: 36. PROFESSOR: Yeah. 1 over 36 for each of them by symmetry.
And by the physics of the dice. So, there is a problem with this. We're computer scientists, not statisticians
in this department. And so we need to worry about things like representing all of this.
So, the size of the distribution, the number of outcomes in the outcome space, the number of possible worlds in the set big omega,
how does that depend on the number of variables and the size of their ranges? So in that example, we had two dice, each
with six possible outcomes. And so that was 6 squared. So this ends up being--
wrong way-- d to the n. Very, very bad. Quickly gets out of hand.
So we can't write out giant joint distributions by hand unless they're very small. Imagine the example I said earlier,
the random variable about how many people show up to a lecture on a given day. Well, that would be determined by the probability
that you show up and you show up and you show up and you show up, every single one of you. And so it's either yes or no for each one of you.
And so that's 2 to the power of a few hundred. You couldn't possibly write down-- and maybe some are correlated.
Maybe you have a friend in this class, you'll usually-- you're more likely to show up if your friend is here. But trying to write down every--
a probability assignment to every single possible outcome would be not remotely feasible. Okay.
But we're going to get to solutions to that later. Now we're still going to talk about understanding joint distributions first in all their--
in all their wasteful specificity. Okay. So, recall that the probability of an event
is the sum of the probabilities of the worlds or outcomes in which that event is true. So an event is a set of worlds.
It is a subset of the set of possible worlds. So we look at all of the worlds, little omega, that belong to our event.
Maybe it's surprising to see the symbol like belongs to when we're talking about events, but an event is just a set of outcomes.
So we sum over all the outcomes that belong to our event of the probability associated with each of these outcomes.
I've explained variations of this a couple of times now, and that's the probability of our event. So, given a joint distribution over all variables,
we can compute any event probability. For example, what is the probability that it is hot and sunny?
Who can tell me? I think someone-- yeah, it's easy. Okay.
0.45. Probability that it is hot? One tiny step harder.
0.6, is that what you said? Oh, no. I assumed you'd gotten it right and did it myself wrong
because I was looking at sunny. Yeah, 0.5. It's the sum of all these in the hot column.
Probability that is that it's hot or not foggy. You can shout it out if you-- once you do the addition, or if you're clever, the subtraction.
0.7-- 0.73. If this is the only-- this is-- whoops.
This is the only thing it has to not be, so it's 1 minus that, 0.73. Okay.
I have a game for you guys. And I need five volunteers to come up. Raise your hand if you want to volunteer.
Yeah, come on. Come on. Yep.
Yep. One more. One more.
Yeah, come on up. All right. Well done volunteering.
I didn't tell you guys you all have the opportunity to get a dollar each. So that's a lesson to the rest of you.
I'm going to show you two options of events, and you get to pick one event. And then I'll randomly sample an outcome,
and if your event happens, you get a dollar. So my advice to you, pick the event with the higher probability.
Option 1 is going to be in red, option 2 in black-- we'll go down the line this way, so you get the first one. But I'll give you--
I'll show you them all so you can think ahead. Okay. Round 1, do you want the event sunny or hot?
STUDENT: I prefer hot. PROFESSOR: You prefer hot. I hear some mumbling out there.
You can reconsider if you like. Here they are again. STUDENT: I'll take sunny.
PROFESSOR: You'll take sunny. Okay, good choice. Let's sample from this distribution.
You had it either way, but yes, it is sunny. It was hot and sunny, in fact, there you go. There's a dollar.
[APPLAUSE] Well done. You can sit down.
Okay, next up. Okay. So recall here I added this in just
to make sure we were all understanding. We're summing up all the outcomes for which the random variable returns sunny-- the weather
random variable returns sunny. So that's these two, 0.6. The other way would have just would have been 0.5.
Okay. Do you want cold and foggy or do you want rainy? STUDENT: Cold and foggy.
PROFESSOR: Okay. Let's see. Sorry.
Bad luck. You didn't get as good options as in the previous one. All right, next up.
Okay, this one's a little-- I'll have to give you a little more explanation. If it's rainy-- and I'll keep sampling until it is.
So conditional on it being rainy, then you win if it's cold. Or I just do one sample and you win if it's sunny.
So do you want if it's rainy, then it's cold, or do you want sunny? STUDENT: And you keep sampling?
PROFESSOR: I'll keep sampling, I'll keep going to this and sampling new spots until we get an outcome that's rainy. STUDENT: What?
PROFESSOR: No. I'll keep sampling until it's rainy. And then, when we get a sample that's rainy,
you win if it's cold. And you lose if it's hot. STUDENT: [INDISCERNIBLE]
PROFESSOR: Okay. So we're going to keep sampling until we land in rainy. Not there, not there, not there.
Nope. Come on. Keep going.
There we go. And it was cold. You win.
[APPLAUSE] All right, here you go. STUDENT: Thank you.
PROFESSOR: You're welcome. Okay. Next up.
Same thing as before. So that was-- so we write that as cold given rainy. So given that it's rainy that was the event that it's cold.
And now we have two of these conditionals. Do you want given that it's cold-- and I'll keep sampling until it is-- it's foggy,
or given that it's not rainy, it's hot? STUDENT: [INDISCERNIBLE] PROFESSOR: Uh huh.
Exactly. STUDENT: [INDISCERNIBLE] PROFESSOR: Yes.
If I've done the math-- I forgot it, but that looks right to me. So, okay.
Given that it's cold, it will be foggy. Not cold, cold and foggy. There you go.
[APPLAUSE] Okay. Last up.
So, this first one, either you get two draws and it has to be-- sorry, this is not going to be as likely as some of the others. Either it has to be hot on the first draw and sunny
on the second draw, or, if you like, I'll just do one draw, hot and sunny, if you-- the second one.
Okay. Yes. I agree.
We're looking for hot and sunny. Nope. Sorry.
[LAUGHS] All right. Thanks to all the volunteers. We're going to have one more-- we're going to have
an opportunity for four more volunteers later in the-- a few slides in the future. Okay.
So, part of the point of that is to try to convince you that a lot of what I'm about to describe is actually very intuitive even though it will look
like a lot of symbols appearing on the screen that you manipulate. It is-- the reasoning that we'll be doing symbolically here
is that you could intuit if money was on the line. Okay. So, the marginal distributions you compute by summing them out.
Literally by writing them on the margins. So, the probability-- if we have the joint probabilities, the probability that the weather is--
probability of what the weather is going to be the probability what the temperature is going to be jointly, then if we want to know the probability that it's sunny,
we just sum over all the y's-- all the temperatures in this case-- to get the probability that it's sunny.
So here, we write in the margins of this table, the marginal probabilities that we get by summing up the rows or by summing up the columns.
And when you're standing right there, you're not going to be thinking about memorizing this formula, you just see that, okay, yeah, obviously
that's the thing that you should do. Any questions about that, how we get marginal probabilities from joint probabilities?
Great. Next thing is conditional probabilities. And again, it just makes sense that you divide here.
In fact, this is how it's defined. So this isn't an insight, this is a definition, but it should make sense.
So there is a simple relation between joint probabilities and conditional probabilities. The probability of a given b is the probability
of b and divided by the probability of b. So if we were looking at these areas here, the probability that it's cold given that it's rainy--
so the rainy area is this brown area plus this orange area. The probability of cold and rainy as this brown area. So we divide this by the sum of them,
this is the total rainy area, and we get the probability of cold given rainy. Or we can see here that the probability of a given b--
we're zooming in from this whole square to just this red circle. The probability of a given b is now how much of this area is taken up by this region.
How much of the area in this circle is taken up by this region. So that's the probability of a and b
divided by the probability of b. So if we want to know the probability of foggy given cold, we take the probability of foggy and cold
and we divided by the probability of cold. How do we get the probability of cold? That's a marginal probability, so we
sum up over all the different ways that it could be cold. And that's what this goes into here. So here's the probability that the weather is sunny,
the temperature is cold. We take that quotient as-- okay, sunny.
So now we'll see sunny given cold rather than foggy given cold. And as I mentioned, this involves summing up--
it's a marginal probability, summing up all the values that don't appear in here. So summing up all the different values of weather.
And we get 0.5. So it's 0.15 divided by 0.5, which is 0.3. Questions about calculating conditional probabilities?
Okay. Because we're building-- we're going to build on this fast. All right, there it is.
0.3. So let's talk about conditioning. So before, we just mentioned conditional probabilities.
So this is just the conditional probability of sun given cold, and now we'll talk about conditional distributions. So the distribution over weather given the temperature,
let's say. So here's the distribution over weather given that the temperature is hot.
Can anyone see the connection between the hot column in our joint distribution and this new distribution over weather given that it's hot?
Yep? Multiply by 2. Any idea why we multiplied it by 2?
Yeah? Right. Exactly.
So these summed up to 0.5, but probability distributions have to sum to 1. Something's going to happen.
And so then we divide by 0.5, which is multiplying by 2, and we get the probability of sun, rain, fog, meteor given hot.
Same for probability of our various weather events given that it's cold. And here's a bit of extra notation.
So yeah, just take a look at this, which I've been breezing over because it's a little boring, but you do need to know it.
This is the notation for this. It's the probability of the weather, and so that's not a single outcome.
So this is going to be a table, not a number. Our weather distribution given the temperature being hot, our weather distribution given the temperature being cold.
And here, this weird set of tables is the probability of the weather given the temperature. So this notation should mean this, a separate probability
distribution for each value that the temperature variable can take. But they're separate distributions.
This is different from the joint distribution. You can tell this is not a joint distribution because they don't sum to 1.
And we're getting ahead of ourselves, but the way that you can convert a fragment of a joint distribution into a new conditional distribution
is basically just by normalizing it. You take the values, you multiply or divide depending on how you're thinking about things to make sure
that it becomes normal, that it is restored to the normal condition of summing to 1. So the procedure, we multiply each entry by alpha.
Alpha is 1 over the sum of our entries to make sure-- and this is analogous to dividing by the total area that was covered by rainy if we were interested in cold given rainy.
But this is the math for it. Okay. So probability of weather given temperature being cold,
we take these numbers, we figure out what they sum to. They sum to 1/2. And we multiply by 1 over 1/2.
Or divide by 1/2, whatever you like. There we go. Okay.
Here is a very basic observation, but one that will come in handy. Take the definition of conditional probability
and move the denominator over there and you get this nice fact. So just take a minute to process that.
The joint-- the probability of a given b is the probability of a and b Divided by the probability of b summing over all the possible a's that might
be kind of hidden in there. And so that's going to be equivalent to saying the probability of a and b is, okay,
start with the probability of b and then multiply the probability of a given b. So here's an example of going from the joint distributions
and-- sorry. The conditional distributions and a distribution over just--
over the variable that's being conditioned on and turning that into a joint distribution. So we could go the other way.
All the information here is in here. We can recover it by looking at the left-hand side and reproducing the right-hand side.
So we have the probability of hot, the probability of cold. We have the probability of rainy given that it's hot. And then we multiply those together,
0.04 times 0.5 gets us 0.02. So we multiply each entry in this times this to get this column.
Same for that column, we multiply by this. Questions about that? Okay.
Ready for another game? I need four volunteers. You can get dollars.
Yeah. Come on up. Three more.
Yep. Sorry if I'm missing hands that are in the back. Yeah, come on up.
Yeah. Okay. You can do the first round.
Well-- okay. So here's the procedure. It's a little different this time.
But it should be intuitive to you that instead of sampling it this way, we're going to sample it this way,
but it's going to be the same thing. So first, we're going to sample the temperature. We're going to sample whether it's hot or cold.
Then we're going to sample the weather from the appropriate-- do you guys want to come over here? Then we'll sample the weather from the appropriate conditional
distribution. So if I sample cold first, I'll sample the weather from this distribution.
If I sample hot first, I'll sample it from this distribution. This is slightly different, actually, than the distribution
before because I've changed these probabilities. These are actually the same. Okay.
Option 1 is going to be in red, option 2 is going to be in black. Okay, so you're-- so--
yeah, will be 5, 6, 7, and 8. Sorry, 6, 7, 8, and 9. So you can look ahead start thinking about which option
you're going to pick, but you first. Do you want the event sunny and hot-- so first, I'll sample hot, and then I'll go here and I'll sample sunny given hot.
Or do you want foggy and cold? STUDENT: Foggy and cold. PROFESSOR: Foggy and cold.
Good choice. Oops, I have to-- okay.
So I'm not doing this one anymore. Let's get rid of that. There are new distributions-- okay, sampling a temperature.
It was cold, good. A good start. Now we sample from our cold weather
conditional distribution. It was foggy. Well done.
[APPLAUSE] Okay. Next up.
Sunny or foggy? STUDENT: Sunny. PROFESSOR: Okay.
Let's go here first. So it's hot. Oh, not necessarily the best start for you.
Sunny. Wait, do you want-- you wanted sunny. Sorry, I thought you wanted foggy Okay
so that was a very good start for you, and you got it. Okay. Next up-- okay, these are when it gets tricky.
I'm going to keep doing this, and it'll and it'll take going through both of them until we get an outcome that is sunny.
And then you win if it's hot if you choose this one or you win if it's cold if you choose this one. STUDENT: I'm going to go with sunny.
PROFESSOR: Okay. How much do you care? STUDENT: It's only a dollar, so--
PROFESSOR: It's-- [LAUGHS] In fact, it would be-- the odds are the same for you either way. Okay, so we're going to go with--
okay, so, sorry, what was it? STUDENT: Sunny. So I want it to be--
I care about it being sunny and not cold? PROFESSOR: So-- okay, wait. So it's given that it's sunny.
You chose hot, right? STUDENT: I wanted hot. PROFESSOR: Okay.
Given that it's sunny, you want it to be hot. Okay, cold. If it's sunny, you're dead.
Nope. Okay, so we go back. Let's try again.
Hot. You got it. STUDENT: All right.
PROFESSOR: So the first time it was sunny, it came from the hot option. No, not a 5.
Okay. Last one. We're going to keep going until it's not foggy.
You win if hot got us there, you lose if you choose this. You win if hot got us there, or you choose this and then you win if cold got us there.
STUDENT: Cold and not foggy. PROFESSOR: Cold given not foggy. Okay, yes, I think that is the right move.
Okay. Hot-- uh oh. Here's our hot weather.
Are we going to get not foggy? Bad luck. All right.
Thanks, guys. So let's just think about that for a second. So the probability of--
actually, let's look at this one. I told you they're the same. The reason is that--
so the probability of-- the probability of hot and sunny is 0.225. The probability of cold and sunny is also 0.225.
You can see that this is 3 times bigger than this, but this is 3 times bigger than that. And so the total probability that it's sunny
is going to be 0.45. And so 0.225 divided by 0.45 is going to be 50/50 for each of these.
So the probability of hot given sunny is 50%, as is the probability of cold given sunny. And just note that that might be a little surprising
because think about reality, sunniness is certainly more typical of hot weather. And sunny given hot is much bigger than sunny given cold.
But since cold is so much more likely to begin with, even though sunniness is more typical when there's heat, these conditional probabilities end up being the same.
And we'll go through the math of how to calculate this written out, but just note that these prior probabilities-- that's what they're called.
These prior probabilities matter, too. The fact that it is so much more likely to calculate cold to begin with-- to sample cold
to begin with certainly weighs in favor of this option. Okay. Ah.
There's the attendance code, and you can look at these various options while you're doing it. Any questions?
Yeah? STUDENT: [INDISCERNIBLE] PROFESSOR: Uh huh.
STUDENT: How did you get to 50%? PROFESSOR: So I-- so you got-- you were with me up to 0.225, right?
Okay. So it was 0.225 for sunny and hot because we did probability of hot times
probability of sunny given hot. So sunny and hot was 0.225. Cold and hot was 0.225.
Those are the only-- sorry. Cold and sunny was 0.225.
Those are all the ways of getting sunny. If it's going to be sunny, it's either going to be hot and sunny or cold and sunny.
Right? So then the total probability of sunny is summing, doing the marginalization thing,
we sum over both of all the ways that it could be sunny. And so we sum the probability associated with each. So that's 0.225 plus 0.225, we get 0.45.
So that's our total probability of sunny. And then the conditional probability is the probability of hot and sunny
divided by the probability of sunny. So that was 0.225 divided by 0.45, which is 50%. For those of you worrying that that's all words
and you're not looking at something associated with it, you'll see symbols associated with all that coming up shortly. Mm-hmm?
STUDENT: You're saying I won a coin flip. PROFESSOR: You won a coin flip. That is right.
Okay. So let's generalize a bit from this fact up here on the top-left.
A joint distribution-- so that was just two random variables. A joint distribution can be written as a product of conditional distributions
by repeated application of the product rule. So, let's say we have three random variables. The joint probability, probability of x1, x2, x3
is equal to this. First, we treat these as one thing and we say it's equal to the probability of x1 and x2 times
the probability of x3 given x1 and x2. And then we do this again. We say that this bit is equal to the probability of x2 given
x1 times the probability of x1. So we just repeatedly apply it. And so in general, that looks like this.
The probability of x1, x2, x3 up to xn is equal to this product of conditional probabilities. The probability of xi given all the ones that came before it.
For x1, this is going to be empty. So one of these is not going to be a conditional probability, it's just going to be a pure probability of x1,
but that's going to be the only one that's not a conditional probability. All the others are going to be conditioned
on the random variables that came before it. And this is true for any ordering, so we can pick a convenient ordering if one is
more convenient than another. Questions about this? How comfortable do you guys feel with conditional distributions
or conditional probabilities? Okay. Yep?
STUDENT: So for the probability [INDISCERNIBLE] PROFESSOR: In the expanded form. What do you mean by expanded form?
Like as a quotient? Like-- STUDENT: As a quotient.
PROFESSOR: Yeah. Right. Okay, good.
So that would be the probability of x1, x2, x3 divided by the probability of x1, x2. So these are the only worlds we're considering, really,
the worlds where x1 and x2 are taking those values. And so when we're looking at the probability associated to worlds where they're all true, we only care
about what fraction of the worlds relative to the worlds that we're considering, which is just the worlds where x1 and x2 are holding.
I don't know, maybe that was more confusing. Yeah. You can see by moving this over to the left-hand side,
this equals this divided by that. Okay. Yeah?
STUDENT: [INDISCERNIBLE] PROFESSOR: Yes, exactly. The left-hand side is the joint probability for all these three
things. So you can think of this as the probability that capital X1 takes the value little x1 and capital
X2 takes the value little x2 and capital X3 takes the value of little x3. Maybe these are dice rolls, for instance.
Or maybe this is the season and this is the weather and this is the temperature. Yeah.
And then on the right-hand side, it's just another way of computing that. It's using conditional probabilities to compute it.
Okay. So when we're doing probabilistic inference in the real world and not just placing bets on things,
we-- typically, there are some variable where we're interested in the probability of it, and there are other variables where
we have evidence about them. So for example, the probability that we get to the airport on time given that there are no accidents,
maybe that's 90%. And these represent the agent's beliefs given the evidence, the agent's uncertainty.
And obviously, if you think about it, probabilities change when you get new evidence. So let's say we add in the evidence.
You look at your clock and you see that it's 5:00 in the morning. Probably you would have already known that,
but let's just say this is a new piece of knowledge for you. Probability that you get to the airport on time given that there's no accidents and it's 5:00 in the morning
is maybe going to be a little higher because there aren't so many people on the road. So maybe that's 95%.
So you add in New evidence and your probabilities about an outcome of interest might change. Adding in evidence is basically conditioning on things.
So when we were doing the game before and I would say, okay, we're going to keep sampling until it's cold and then you win if it's hot or whatever it was, no.
We keep sampling until it's cold and you win if it's sunny, let's say. That's basically me telling you, okay, for this sample,
the outcome is going to be-- it's going to be some kind of cold outcome. Maybe a sunny cold, maybe a rainy cold, but it's
basically like we're getting new information. When we get information, then we can condition on it. Obviously we're only going to be interested in outcomes
that are consistent with the information we've gathered. Here, we might add some more evidence maybe that it's raining, and now the probability
that we get to the airport on time might go down. Maybe conditioned on this new evidence, it's going to be lower.
So it could go up, it could go down. It should do both, generally, over time. So observing new evidence causes beliefs to be updated.
Cool, great. How do we do inference? Okay.
The first way I'm going to show you is a very inefficient way, but the most-- it just follows the most straightforwardly
from the definitions I've given you so far. So let's say we're given the joint distribution. Even though this could be massive in many cases,
in many settings, there's no way we could possibly have access to this. Let's just say in a theoretical nice world,
we have access to the whole joint distribution. For every set of values that our variables might take, we can look up the probability that those variables
take those values. And let's-- just for our own purposes talking about this, let's say that some of these variables are evidence
variables. These are things where we'll see the outcome. Some of them are query variables.
These are variables where we're interested in the probabilities that they're going to take these values but we can't see it directly.
And maybe there are some hidden variables that will be relevant to the relationships between the evidence and the queries, but we don't see them
and we're not really interested in talking about them either. Okay. So we're going to want, often, a probability distribution
over our query variables, and this is in bold to indicate that it's representing a set of random variables.
So this might be like x2, x7, x12. So we want a probability distribution over all of the values that our query variables
might take given all the values that our evidence variables have taken-- that we've seen them take. So we saw that it was rainy and we
want to know the probability that we get to the airport late. So step 1 is we select the entries in the table consistent with the evidence.
Step 2, we sum out all of the hidden variables from the model to get a joint of query-- of the query probability-- sorry,
the probability is associated with the queries-- query random variables and the evidence random variables. But it's not truly a joint probability
because it won't be normalized. Step 3, we have to normalize. So, here, we're summing out all of the H's, and I'll
show you an example in a second to only get the query and the evidence. And then since we'll have crossed out a lot of outcomes,
since they weren't consistent with the evidence, the probabilities will sum to less than 1, we'll have to divide by the total probability
to get a good-- to get a true probability distribution. So that's that that normalizing factor alpha. Okay, here's an example.
So we've added season into the mix here. And now we want to know the probability distribution over seasons given that it's sunny.
So first step, we enumerate all the options where it's sunny and we cross out the rest. So already we don't have a probability distribution
because these sum to less than 1, but you can think of it as of a partial probability distribution-- these are all the outcomes consistent with the sun
that we've observed. Questions about this step? Next step, we sum out the irrelevant variables.
So we don't care whether it's hot or cold, we just care about whether it's summer or winter. So if we're interested in the probability that it's summer,
we have to sum both of these entries because there are two ways of getting to the summer outcome, through the hot--
the hot option, the cold option. So we sum these together, we get 0.45. We sum these together, we get 0.25.
And we're not done. I mean, it's not the probability of summer is 0.45 and the probability of winter is 0.25
because those are the only options. So we have to normalize. So the ratio of probabilities is correct,
but the absolute probabilities are not until we divide by the sum and then we guarantee that the probabilities will sum to 1.
I haven't done the math. This equals some number. I don't think it's so instructive to tell you
what it is. This is how you calculate it. Okay.
Yeah? STUDENT: Temperature? PROFESSOR: Temperature-- right.
So this was-- in this example, temperature was one of the hidden variables. We weren't observing it, we didn't care about it,
so we just sum over it. STUDENT: [INDISCERNIBLE] PROFESSOR: So-- yeah, basically.
So we had to sum-- for any value of interest, we had to sum over all options of temperature.
So for the summer option of interest, we sum over all values of temperature that could be associated with the summer option
because each of those is going to have its own entry in the table of joint probabilities. Yep?
Normalize is just make it sum to 1. So this is a bad probability distribution. If I told you that I was going to flip a coin
and it was going to come up heads 45% of the time and tails 25% of the time, you would probably tell me like, are you missing something?
Like, that-- might anything else happen? So if the answer is no, if those are the only options, then we had better divide by the sum
to make sure that they're proper probabilities. In the game from before, when we were conditioning on it being a--
when we were conditioning on it being rainy, if we're going to wait until we get a sample somewhere in here, then we have to expand everything because this is only
a small part. So we have to divide by the total-- by the probability of what we're conditioning on.
And in general, I'd recommend, if you forget like, okay, what's the step that comes next? Because if you're just treating this as steps
to memorize or just all procedural stuff with manipulating symbols, go back to the example of you're playing a game, you win in this setting,
you lose in this setting, I think it's a much more intuitive setting and I think it'll help you know what to do much more reliably.
Okay. So, some problems with inference by enumeration. The worst-case time complexity is
exponential in the hidden variables because we're summing over all the possible values of this hidden variable and all the possible values
of this hidden variable and this one and this one, and all the different combinations of values they can take is exponential.
Worse still is the expectation that we're going to be able to store the joint probability distribution. I could not hope to store the joint probability
distribution of showing up, you showing up, you showing up-- every single one of you showing up all together to the 100-something--
to the couple hundred, there's just no way to list all those possibilities. And worse than that, if we're going to get good,
often the way that we'll get some idea about the probability of an outcome is by seeing it happen. So if we see that something happens 10 times out
of 100 opportunities to happen, that's decent evidence that it's going to happen-- that maybe the probability was about 10%.
If we have exponentially many entries in a table-- if we have a table of exponential size-- so we have exponentially many entries
that we're going to need to fill in with data, it will cost-- then we need exponentially many data points, and getting data in the real world takes money.
So that's just-- there's no way we could fill in this table feasibly. Even if we could write it down, we
wouldn't know what numbers to write. So, we need to do better things. And what we're going to build up to is Bayes' nets.
But to get there-- and we're not going to get to Bayes' nets today, but what we can get to today is Bayes' rule.
So, let's talk about that. It used to be that if you rearranged a definition, they would name a rule after you.
Now it takes a little more work to get your name on something. But look, here-- so the definition of conditional probability was basically
probability of a given b was the probability of a and b divided by the probability of b. Okay, so we've moved the probability of b over there.
And now we get the product rule. We can write the product rule both ways. And then we divide by this and we get this.
That's Bayes' rule. Just follows from the definition of conditional probability. That is a picture that may or may not
be of the Reverend Thomas Bayes. And that's his rule. Okay.
So why-- I mean, why do-- maybe you've heard people making huge fusses about Bayes' rule and how amazing it is.
I mean, it's just so straightforward manipulation-- such a straightforward manipulation of symbols from a definition, why is this useful?
Okay. The reason why it's useful is that often, it'll be easy to understand the conditional probability in one
direction and hard to understand it in the other direction. So this allows us to get the direction where it's hard to come up with it.
And I'll give you an example that should make it more clear. The other way of understanding it is we start out with a prior probability of a belief
that a is true. And then if we learn b, then we condition on b. So we get evidence that b has taken a certain value.
So now we want to see how that's changes our belief in a. And this is the formula for how we-- for how you do it. How you start with one belief about the probability of a
and modify it into another as new evidence comes in. So here's an example where it might be relatively easier to get this stuff on the right, but we care
about this thing on the left. So let's say d is our data, h is our hypothesis. Then just-- then you can write this
as the probability of a hypothesis given the data is equal to all this stuff. This was your prior belief in the hypothesis
before you saw anything. So maybe that's just a measure of how a priori likely you think it is, how much sense it would make
before you've seen evidence. Any hypothesis should come with its idea of how likely it is that you're going to see different data.
And then this probability of the data also depends on these-- you can sum over all the possible hypotheses that you're considering and this will give you--
so this will-- going from here to here would be marginalization. So yeah, we're basically doing marginalization.
We don't necessarily have this one, but if we have all the things in the row or all the things in the column or whatever,
then we can sum them together to get this. So let me give you an even more specific example. Let's say that you have a fever.
That's your data, so you know that's true. What you don't know is whether you have the flu or you have COVID.
So your hypotheses are, hypothesis 1, I have COVID; hypothesis 2, I have the flu. And you can't look up online what's
the probability I have COVID given that I have a flu because it changes wildly depending on what's going on in the world.
For instance, 10 years ago the probability that you had COVID given that you have a flu is 0 because COVID didn't exist 10 years ago.
Maybe in different countries it's different. But what's very consistent is if you have a fever, what's the probability you get a flu?
If you have COVID, what's the probability you get a flu? So those conditional probabilities, the conditional of the data given the hypothesis, those
are consistent. If we get lots of data in one country, we can generalize to our country.
So we can work with these conditional probabilities. Or if it were a less complex system than biology, if we're talking about some physical system,
maybe we could derive it computationally. Maybe the hypothesis would just give us a function that would tell us the probability
of different data. So in any case, it is often much easier to understand, to learn from, to generalize facts--
yeah, to generalize our understanding of conditionals in this form, probability of data given a hypothesis, than it is in here.
But this might be the thing we're interested in. I mean, I have a-- I have a-- sorry, I have a fever and I want to know,
do I have COVID or do I have the flu? And so it's a common setting where the thing we're interested in is not the thing that
is easiest to learn about. So here is a way to get the thing of interest using the things that you can get cheaply.
Here it is written in another way. Easy to learn about, know about the probability of an effect given a cause, but often
we'll want to infer what a cause must have been given observed effects. Here's another example in addition to the fever
and illness example. Here is stiff neck given meningitis. So let's say these are our probabilities.
Let's say the prior probability of meningitis is 1 in 10,000. It's not very common for people to get meningitis. And so 1 in 10,000 people have meningitis, 1 in 100 people
have a stiff neck. These are things we can look up from all the data, from people visiting the doctor.
And we know something about how meningitis works. We can collect data about people who get meningitis. And so maybe we can confidently say the probability
that you have a stiff neck, if you have meningitis, is 80%. But I have a stiff neck and I want to know if I-- like, if that means I should start worrying about meningitis.
Here you go. Here's how you compute it. So the probability of meningitis given a stiff neck
is computed according to that formula. We plug in all the numbers and we get 0.008. So even though a stiff neck is very typical for meningitis,
meningitis isn't very typical for a stiff neck because there are so many stiff necks out there and meningitis is pretty rare.
So, it's very small, but it's 80 times bigger than it was. So the priors matter. This is going to be proportional to your prior belief as well.
So it's not just-- there's a common kind of bias that people have, which is they think about typical--
they think about explanations such that what they're seeing would be typical for that explanation, but ignore the base rates.
They ignore how likely that explanation is to begin with. So there's a famous example where it's like-- let's say, what's the probability that a banker is
a woman versus what's the probability that a banker is a woman who really cares about her career and is trying to slough off these ideas about stereotypes
of people and maybe informed by some stereotypes that they had, thought that the latter probability was bigger than the former probability even though the former probability
can-- the former event contains the latter event? So if everything I told you about the second case
was true in the first case as well, so the first probability had to be higher, but the second one, it sounded more typical to people.
And so they ignored the base rates, and so gave answers that, relatively to each other, didn't make any sense. So you have to keep--
you have to keep-- you have to pay attention to the prior probability, too. So it's still very small, but it's
80 times bigger than it was. That's the sort of thing evidence can do. That's not to say you shouldn't get your stiff neck checked out
because it would be very bad if you did have meningitis. It could lead to all sorts of complications, maybe paralysis, I don't know.
Don't quote me on that, I'm not a doctor. So, it's still good reason to get it checked out. But, all told, until you get other tests,
it's still probably not meningitis. Okay. I'm going to introduce one more topic,
and then the rest will be left for Bayes' nets-- the discussion of Bayes' nets starting next week, and that topic is independence.
So what is independent mean? Two variables, x and y, are independent if the joint probability--
for any two outcomes, the joint probability is equal to the product of the probabilities between them. So these are factors.
The joint distribution, it factors into two simpler distributions. Another way to write it that's equivalent,
given the definition of conditional probability or the product rule, we can use the product rule like this-- so we substitute for this.
We cross off P of y on both sides. So this is the same as saying P of x given y is the same as P of x.
And that's the same as saying the probability of x happening-- your belief about the probability of x should not change if you observe y.
So your belief about what the result of the second time I roll a die is shouldn't change when you observe what it was the first time I rolled a die.
So you probably are-- if I was about to roll a dice-- or roll a die twice, you're probably expecting the second time
I roll it, 1-in-6 chance for each of the outcomes. And if I roll the first one and the first one comes up 4, it shouldn't change your beliefs about what it's
going to be for the second one. That's another way of saying the two dice rolls are independent. So there's our example.
The probability that roll 1 is a 5, roll 2 is a 3 is the same as the product of the probability that the first roll is a 5 with the probability
that the second roll is a 3. So 1 in 6 times 1 in 6 is 1 in 36. And, if you have things that are independent,
you can much more concisely characterize the joint distribution because you don't need to write down a big table,
you just write down each one individually. So if I have n coin flips, I can just-- and they're all independent, I can say,
okay, each one is 50 over 50. I don't have to write down all 2 to the n entries for all sequences of n coin flips that might occur.
So, independence is incredibly powerful. It allows you to exponentially reduce the size of a state space-- can those of you
who are here for the next class maybe back up or quiet down a little bit because we still have a little more time here? Thanks.
I am-- so, it can allow you to represent your beliefs or your-- yeah, it can allow you to represent your beliefs in much less space.
But unfortunately, it is extremely rare that any two things are actually independent. What is much more common is conditional independence.
And this is going to be the main tool that we use for tractably representing vast amounts of random variables that have some connection to each other.
So as a little teaser for what's going to come next time, here is an example where we have three random variables. Is there traffic or not, is it raining or not?
Guys in the back, can you please be a little quiet? We do have a little more time here. Thank you.
Okay. Is there traffic or not? Is it raining or not?
Is there-- is some person going to have an umbrella or not? Now no pair of these are independent. If you see that someone has an umbrella,
you can guess that there's going to be traffic because, well if they have an umbrella, it's probably raining. If it's raining, there's probably traffic.
And obviously, if you see someone has an umbrella, you can guess that's going to inform you, that's going to change your--
change the probability you assign to it being raining, so those are obviously not independent. You can observe one and change your beliefs about the other.
Same for raining and traffic. You can observe one, change your beliefs about the other. But two of these variables are conditionally independent,
probably, given the third. Conditional on whether it's raining or not, it makes sense that learning about the traffic
is not going to tell you about whether someone's carrying an umbrella or vice versa. If I already know that it's raining
and I see that you're carrying an umbrella, that's not going to give me any more information about whether I can expect traffic
on the Bay Bridge. So conditional on one thing, that will sometimes make it the case that learning about a second thing
doesn't change your beliefs about a third thing. And that's conditional independence. And that is much more common.
It's how we are able to factor our understanding of various things that we're uncertain about and reason tractably.
And that's going to be the basis for the next couple of weeks of what we're talking about, which is tractable inference about uncertain outcomes.
All right. You can come up with any questions if you like. Thanks, guys.
Probability is essential in AI because the real world is full of uncertainty. Classical logic cannot handle partial observability (incomplete knowledge), noisy sensors (imperfect measurements), or model limitations (inability to process all data). Probabilistic models allow AI systems to make rational decisions by quantifying uncertainty and summarizing ignorance.
A possible world is a complete, mutually exclusive outcome (e.g., a die roll result of 3). An event is any subset of possible worlds (e.g., the roll being odd). The probability of an event is the sum of the probabilities of the worlds inside that subset. The probability model assigns a non‑negative number to each possible world such that they sum to 1.
A joint distribution describes the probability of combinations of multiple random variables (e.g., P(Temperature, Weather)). A marginal distribution is the probability of a single variable, obtained by summing over all other variables (marginalization) in the joint. You can derive marginals from the joint, but not vice versa, which is why the joint contains the full probabilistic relationship.
Conditional probability P(A | B) is defined as P(A ∧ B) / P(B), renormalizing the joint distribution to focus only on worlds where B is true. This updates beliefs based on new evidence. The chain rule (P(A, B) = P(A | B) * P(B)) follows from this definition and allows any joint distribution to be factorized into a product of conditional distributions.
Bayes’ rule (P(A | B) = P(B | A) * P(A) / P(B)) allows us to compute the probability of a cause given an effect, even when we only have data about the effect given the cause. For example, although 80% of meningitis cases cause a stiff neck, the probability of having meningitis given a stiff neck is only 0.8% because the base rate of meningitis is very low. This illustrates the critical role of prior probabilities.
Two variables are independent if P(X, Y) = P(X) * P(Y) for all values – knowing one never changes belief about the other. Independence drastically reduces the state space (e.g., n independent coin flips need n probabilities instead of 2^n). Conditional independence is even more powerful: two variables become independent once a third variable is known. This principle is the foundation of efficient probabilistic models like Bayes’ nets.
Inference by enumeration is a brute‑force method: select all joint entries consistent with the evidence, sum out hidden variables, and normalize. However, it is intractable for real‑world problems because the joint distribution grows exponentially with the number of variables. Bayes’ nets address this by exploiting conditional independence to represent the joint distribution compactly, making inference feasible.
Keep this summary
Save it to LunaNotes and it becomes a real note in your library — editable, searchable, and ready to turn into flashcards or a diagram. Free to start.
Save to LunaNotesOr summarise for another video.
This summary and transcript were automatically generated using AI with the Free YouTube Transcript Summary Tool by LunaNotes.
Related summaries
Bayes Nets: Mastering Probabilistic Graphical Models for Uncertainty
Comprehensive lecture on Bayes Nets covering conditional independence, probabilistic graphical models, and real-world applications from traffic analysis to Ghostbusters ghost tracking. Learn how Bayesian Networks reduce exponential probability distributions to linear complexity through causal structure encoding.
Probability & Statistics: The Ultimate Guide to Modeling Uncertainty (Course Intro)
Professor Steve Brunton launches an exciting new short course on probability and statistics. This introductory overview explains why probability is a foundational tool for data science and machine learning, provides real-world examples from thermodynamics to weather, and outlines what you will learn in the first 10 hours dedicated to probability.
Calculating Conditional Probabilities Using Tree Diagrams and Dice Rolls
This summary explains how to calculate probabilities and conditional probabilities using tree diagrams, including examples with bus arrival times, ball selections without replacement, and dice rolls. Key concepts such as intersections, complements, and conditional probability formulas are demonstrated with step-by-step calculations.
Introduction to Probability and Statistics: Key Concepts and Terminology
In this video, Dr. Gajendra Purohit introduces the fundamentals of probability and statistics, covering essential terminology, types of events, and key concepts such as random experiments, sample space, and probability calculations. The session aims to provide a solid foundation for students preparing for advanced mathematics exams.
Comprehensive Introduction to AI: History, Models, and Optimization Techniques
This lecture provides a detailed overview of Artificial Intelligence, covering its historical evolution, core paradigms like modeling, inference, and learning, and foundational optimization methods such as dynamic programming and gradient descent. It also discusses AI's societal impacts, challenges, and course logistics for Stanford's CS221.
Most viewed summaries
A Comprehensive Guide to Using Stable Diffusion Forge UI
Explore the Stable Diffusion Forge UI, customizable settings, models, and more to enhance your image generation experience.
Kolonyalismo at Imperyalismo: Ang Kasaysayan ng Pagsakop sa Pilipinas
Tuklasin ang kasaysayan ng kolonyalismo at imperyalismo sa Pilipinas sa pamamagitan ni Ferdinand Magellan.
Mastering Inpainting with Stable Diffusion: Fix Mistakes and Enhance Your Images
Learn to fix mistakes and enhance images with Stable Diffusion's inpainting features effectively.
Pamamaraan at Patakarang Kolonyal ng mga Espanyol sa Pilipinas
Tuklasin ang mga pamamaraan at patakaran ng mga Espanyol sa Pilipinas, at ang epekto nito sa mga Pilipino.
How to Install and Configure Forge: A New Stable Diffusion Web UI
Learn to install and configure the new Forge web UI for Stable Diffusion, with tips on models and settings.
Found this summary useful?
Take it with you. One click puts it in your own LunaNotes library.
Save to LunaNotes