Understanding Bayes Nets: A Comprehensive Guide to Probabilistic Graphical Models
Video Summary: Bayes Nets & Conditional Independence
This lecture introduces Bayes Nets (Bayesian Networks) as a powerful framework for modeling uncertainty in complex real-world scenarios. It builds on fundamental probability concepts; for a refresher, see Introduction to Probability and Statistics: Key Concepts and Terminology. The professor demonstrates how to move from basic probability concepts to sophisticated graphical models that capture conditional independence relationships, dramatically reducing the computational complexity of joint probability distributions from exponential to linear scale.
Key Learning Objectives
- Master the transition from strict independence to conditional independence
- Understand how Bayes Nets encode causal relationships
- Learn to construct and interpret directed acyclic graphs (DAGs)
- Apply Bayes Nets to real-world problems (Ghostbusters, insurance, car troubleshooting)
Keywords
Bayesian Networks, conditional independence, probabilistic graphical models, Bayes Rule, joint distribution, causal inference, Naive Bayes, directed acyclic graph
Core Concepts
1. Probability Foundation Review
Basic Rules of Probability
- Events & Probabilities: Each event has probability between 0-1, summing to 1 across all possible worlds
- Random Variables (X): Functions mapping outcomes to values
- Marginal Distribution: Summing joint distribution over one variable: P(X) = ΣP(X,Y)
- Conditional Probability: P(X|Y) = P(X,Y)/P(Y)
- Product Rule: P(X,Y) = P(X|Y) × P(Y)
- Chain Rule: P(X1,X2,...,Xn) = Π P(Xi|X1,...,Xi−1)
Independence Types
- Strict Independence: P(X,Y) = P(X)×P(Y) - rarely holds in real world
- Conditional Independence: P(X|Y,Z) = P(X|Z) - more practical and common
2. The Traffic/Rain/Umbrella Example
This three-variable system demonstrates why strict independence fails in real scenarios and how conditional independence provides structure.
| Variable | States | Key Finding | |----------|--------|-------------| | Rain | True/False | 30% probability | | Traffic | True/False | 35% probability when true | | Umbrella | True/False | Depends on rain, not traffic |
Crucial Discovery: Umbrella is conditionally independent of traffic given rain - knowing rain status makes traffic information irrelevant for umbrella prediction.
3. Ghostbusters: Naive Bayes in Action
This interactive example demonstrates probabilistic tracking in a 3×3 grid.
Problem Setup:
- Hidden ghost in 9 possible locations
- Noisy sensor returns colors (red/orange/yellow/green) based on proximity
- Goal: Infer ghost location from sensor readings
Key Insights:
- Without Independence: 9×49 = 2.3 million possible outcomes
- With Conditional Independence: 9×4 = 36 conditional probability tables
- Sensors are conditionally independent given ghost location
This example illustrates a core application of Naive Bayes. For a deeper dive into the probability foundations that power AI models like this, see Introduction to Probability: Foundations for AI & Bayes' Nets.
4. Building Bayes Nets: Technical Framework
Structural Components
- Nodes: Random variables in the system
- Edges: Dependency relationships (arrows = causality or influence)
- Conditional Probability Tables (CPTs) : Quantitative relationships
Size Comparison
| Method | Parameter Count | Formula | |--------|-----------------|---------| | Full Joint Distribution | 31 | 25-1 (for 5 binary variables) | | Bayes Net | 10 | Σ(K×parents) |
5. The Burglary/Earthquake/Alarm Example
This canonical example (by Judea Pearl) illustrates Bayes Net construction and inference.
Network Structure:
Burglary ──┐
├──> Alarm ──> JohnCalls
Earthquake ─┘ └──> MaryCalls
Conditional Probabilities:
- P(Burglary) = 0.001
- P(Earthquake) = 0.002
- P(Alarm|Burglary,Earthquake) = 0.95 (both), 0.94 (burglary only), 0.29 (earthquake only), 0.001 (neither)
- P(JohnCalls|Alarm) = 0.90, P(JohnCalls|¬Alarm) = 0.05
- P(MaryCalls|Alarm) = 0.70, P(MaryCalls|¬Alarm) = 0.01
6. Model Selection: The Accuracy-Efficiency Trade-off
Bayes Nets occupy a spectrum between two extremes:
| Model Type | Accuracy | Efficiency | Use Case | |------------|----------|------------|----------| | Strict Independence | Low | Excellent | Coin flips, dice rolls | | Naive Bayes | Medium | Very Good | Ghostbusters, spam detection | | Sparse Bayes Net | High | Good | Insurance risk assessment | | Full Joint Distribution | Perfect | Poor | Small systems (<6 variables) |
Practical Applications
Insurance Risk Assessment
- Query Variables: Medical cost, liability cost, property cost
- Evidence Variables: Age, driving record, car type, garage status
- Hidden Variables: Risk aversion (inferred from choices)
Car Troubleshooting Diagnostic Network
- Symptoms: Car won't start, lights dim
- Causes: Battery failure, alternator issues, starter problems
- Structure: Multiple causes can produce same symptoms
Advanced Concepts
Causality vs. Correlation
- Arrows represent conditional independence, not necessarily causation
- Example: Rain → Traffic (causal) vs. Traffic → Rain (mathematically equivalent but counterintuitive)
- Choose causal direction for easier probability elicitation
Explaining Away Effect
When multiple causes can produce the same effect, observing one cause reduces the probability of alternative causes:
- P(Earthquake|Alarm=True) decreases when we learn Burglary=True
Technical Requirements
- Graph Type: Directed Acyclic Graph (DAG)
- Node Capacity: D values each
- Parent Limit: K parents per node (sparsity assumption)
- Memory: O(N×D^K) vs. O(D^N) for full joint
Next Steps
- Apply chain rule with Bayes Net structure
- Implement inference algorithms for probability queries
- Extend to larger, more complex domains
Summary
Bayes Nets provide a principled approach to modeling uncertainty by exploiting conditional independence relationships. This transforms intractable exponential probability spaces into manageable linear representations while maintaining model fidelity for practical applications.
PROFESSOR: A couple of announcements before we get started. Let me see if I can turn the volume down here.
Is that okay? Nice. So yeah, Homework 3 is due today,
and then Project 3 will be due next Tuesday. Homework 4 will be out later this week, due next Friday. And just sort of a friendly reminder that two weeks
from today will be the midterm. So there's more info on the website. We'll have review sessions, and we'll
have discussions that will help you prepare for that. So yeah, keep an eye out for the announcements posted on Ed. Today, we're going to talk about Bayes Nets, which
is going to build on the lecture from last Thursday, and essentially what we're going to be doing is we're going to be building these concepts
or building these models out of the concepts we learned in probability so that we can help ourselves answer detailed questions about uncertainty.
So this idea of taking the concepts from probability and leveraging them and making them work for us. So just to recap what you guys talked about with Michael
last time, there's some basic rules in probability, and one of those rules is that there are events that can happen.
And when an event happens, it makes the set of possible worlds, it'll sort of restrict that to whatever actually happened.
Each of these events is associated with a probability. That probability has sort of a likelihood of occurring, and these numbers exist between zero and one.
They have to sum to one, all these classic probability things. We can have events.
We can talk about events. An event is just a subset of the possible worlds that meet a certain condition.
So if we have this Event A, it might be like, I made it to lecture on time, and it could have been done in any number of ways.
I could have taken the bus. I could have walked here. I could have taken the train.
The Event A just means I got here in time, and it doesn't really matter how we did that. But each of the probabilities associated with me getting here
by taking the train, me getting here by walking, me getting here by riding my bike has a probability. And then the total of the probability for all the events,
all the outcomes that are consistent with that event, that's what we're trying to add up here. We'll also talk about random variables.
We give these capital letters like X, and we say that these are functions that take in an outcome and tell you a value.
So each of these values, now, each of the values that a random variable can take on, has some associated probability with it.
And that's going to be dependent on which outcomes have which values. We can also sum out-- if we have a joint distribution for P
of X and Y, we can sum out to get the marginal distribution for X or the marginal distribution for Y. All we have to do is add up all the different ways.
This is, how did I get here? We add up all the different ways that we could see the variable Y manifesting itself,
and then what's left is just X. So that's the probability distribution over X. And the way that we compute that is we just add things up.
We also talked about conditional probability distributions. So this says if I have some information, Y, and I want to figure out some information, X,
I can model the probability distribution of X conditioned on Y by thinking about all of the events where Y is true as sort of like my normalizing constant.
This is like the total probability that's available if Y is true or the total probability that is available if Y is false.
And then I look at all of the events in the joint distribution X and Y, and I take the subset of the ones that
match my sort of XY condition. So this is just the conditional probability from last time. And you can also rearrange this to get this product rule that
says P of X given Y times P of Y is equal to P of XY, the joint. And you can do it in either direction. So P of Y given X times P of X, and this is just Bayes Rule.
And this also allows us to write arbitrarily complex joint distributions in terms of a product of their sort of conditional probabilities.
So P of X1, X2, X3, Xn is just multiply the conditional probability distributions for each one and then condition on all the variables
that came before that. We also talked about independence, and I wrote strict independence here because in a few minutes,
we're going to talk about conditional independence, and we'll contrast the two. But independence is basically two variables
are independent if it's the case that the probability of X and Y is equal to the product of the two marginals, P of X and P of Y. So the joint distribution factors into these two
different marginal distributions, and you can multiply them together. Or we could also write it this way.
We could say that the conditioning information doesn't change what we know. So if I have this conditional distribution P of X given Y,
and I find out that Y is true, that doesn't change my distribution over X. So learning Y doesn't tell me anything interesting about X.
Or learning X doesn't tell me anything interesting about Y also. So all these things are equivalent.
And so an example of this is you roll a six-sided die two times, and the probability that you get a 5 on the first roll doesn't affect the probability that you get a 3
on the second roll. So the two die rolls are independent because of the way that the physics of these dice work.
So one die roll doesn't affect the next one. Yeah, and so, basically, the world doesn't really work like this, in general.
Like, we have to design these materials like dice and coin flips and roulette tables and slot machines in order to make these assumptions
hold, these strict independence assumptions. But when they do hold, it's really nice. We can, instead of taking this joint distribution over n
different coin flips, we can-- which is going to take us exponential space to write down because there's 2 to the n different ways
that all those coin flips can come out. And if we want to write down the joint distribution, we have to write down a probability
for each of those possible combinations of coin flips. Now we can say, all right, I know that these are independent coin flips,
so I'm just going to write down the probability table for each coin flip separately. And I get n different tables, and each of those tables
has two parameters. Really, it has one parameter because they have to sum to one. So now instead of 2 to the n, we're
looking at something that's n. So we've gone from exponential to linear, and that's a huge, huge savings.
But, of course, this only works when we have these nice materials like dice and coin flips. It doesn't work in the real world.
So what happens if we try to apply this idea in the real world? So here's this example that we talked about towards the end
of last time. We have traffic. We have rain.
We could be holding our umbrella or not holding our umbrella, and we want to model this scenario. So we model it with this joint distribution.
The joint distribution specifies the probabilities for each of these outcomes. So each of these variables is binary,
and so we have this table. We can put all the probabilities in here. You can check that those sum to one,
and this is our joint distribution. This is like our contingency table. So if I wanted to compute the marginals, how would I do that?
How would I get the probability? Let's say the probability that rain is false. How could I compute the probability that rain is false?
Yeah. STUDENT: [INDISCERNIBLE] probability. PROFESSOR: Say it again.
STUDENT: The four p values. PROFESSOR: Add up the four p values. Which four?
STUDENT: From 0.503 to 0.414 PROFESSOR: These ones? STUDENT: Yes.
PROFESSOR: Yeah, why are these the values that we want to add up? So you're saying add up all these ones.
These are consistent with we said we want the probability that rain is false. And so all of these have rain is false,
and so the probability of each outcome needs to be added up. And we don't care about any of the other variables. So does anyone want to add those up for me?
0.7? So 0.7. Do we know anything else?
What do we know? Yeah. STUDENT: We need to normalize, right?
PROFESSOR: We need to normalize. What do you think we need to normalize? STUDENT: [INDISCERNIBLE].
PROFESSOR: No, I think we're okay here. We'll have to normalize when we get to conditional probabilities, but right now we're okay.
We can just add these things up. But what else-- what other information do we have about the probability of rain?
Yeah. STUDENT: The probability of rain being true is just 0.3 confidence.
PROFESSOR: Yeah, right. So we also know 0.3. Yeah, so that's great.
So we can actually fill this table in, and we get 0.7, 0.3. So let's do this one more time. Let me ask about the probability that traffic is true.
Which-- what do we need to do now? Yeah. STUDENT: And it's more [INDISCERNIBLE] traffic is true.
PROFESSOR: Yeah, exactly. So we look at our table. We find the values where traffic is true.
Traffic is true here. Traffic is true here, so we add up these values. And when we add those up, what do we get?
STUDENT: 0.35. PROFESSOR: 0.35. That sounds right.
And do we know anything else? STUDENT: So in 0.65-- PROFESSOR: 0.65 for false, yes.
So 0.65, awesome. So we can fill that table in. I'll save you the trouble.
We can do this for umbrella as well. So if we were going to try and model this three variable system with rain, traffic, and umbrella as being strictly
independent, these variables being strictly independent of each other, how could we test that, given what we've already calculated?
Yeah. STUDENT: Multiply it together. PROFESSOR: Multiply them together, yeah.
And see what? STUDENT: See if it equals the p value you have already. PROFESSOR: See if that equals the p value
that we have already, great. So let's do that. Let's do that for the first entry.
So let's look at the probability of a rain being false, traffic being false, umbrella being false. And what numbers do I have to multiply?
STUDENT: [INDISCERNIBLE] PROFESSOR: This one, this one, and this one. And does that come out to 0.504?
Does anyone know what that comes out to? STUDENT: 0.314. PROFESSOR: 0.314.
Well, that doesn't seem like it's 0.504. Should we do one more? Let's do this one.
True, false, true. We can do it in another color. So true, false, true so 0.3 times 0.65 times 0.31.
STUDENT: [INDISCERNIBLE] PROFESSOR: 0.06? So these are not coming out to be the same thing,
so that's interesting. So what that's telling us, if we fill in this whole table, you'll notice that these are not the same thing as we started
with. So we take our joint distribution, we compute the marginal distributions,
we multiply them together, they're not the same. So that tells us that we don't have strict independence here. That's saying that the variables are not
independent from each other. But what do we always know that we can do? We talked about this just now.
We know that this chain rule thing works no matter whether we have independence or not. So we can always write things as a product
of these conditional probability tables, and we're building up the number of conditioning variables in each case.
So rain, we already calculated. We have this probability that rain occurs. And then given that we have rain,
we know that we can write down the probability of traffic conditioned on rain. So now let's practice writing down
these conditional probabilities. So what about this entry right here? This is the probability that we do have
traffic when it's not raining. So what do we have to look at in our table to figure this one out?
It's a little more complicated than before, but not so much more complicated. So we know that rain is false.
So we can look over in our table, and we say, well, conditioned on rain being false, let's start looking at these possible worlds.
And now we want the probability of traffic being true conditioned on being in this blue box. So let's draw that in red.
So the probability that traffic is true is comprised of these two events. So what do we want?
We want basically the amount of probability in this red box divided by the amount of probability in this bigger blue box.
So this is a fraction. This is saying the conditional probability of traffic being true conditioned
on rain being false is this red box divided by this blue box. So what's in the red box? We get 0.14.
And what's in the blue box? 0.7. Is that right?
And so 0.14 divided by 0.7-- let's see. 0.2?
Great. 0.2. Do we know anything else?
What else do we know? Yeah. STUDENT: It's 0.8 that also falls?
PROFESSOR: Yeah, why is that? STUDENT: Because it needs to add up to 100%. PROFESSOR: Yeah, it's 0.8 right here
because this needs to add up to one, exactly. So let's clear these things out. We can fill in this whole table.
We can do the same thing. Let's do it one more with-- now we have-- so essentially, let's recap what we're doing.
We're saying let's keep track of all the events where my conditioning variable has the value that I want it to take, and then let me look at the fraction of those cases
where the particular condition on traffic held. So we looked at this one first. Now let's look at, let's say, this one.
So this is the probability that I don't have my umbrella when it is raining but there's no traffic.
So what is that? Which numbers do we have to look at in this table? Let's number them 1, 2, 3, 4, 5, 6, 7, 8.
Yeah, who thinks they know? Yeah. STUDENT: 5 and 6.
PROFESSOR: 5 and 6. So 5 and 6 are these two rows. And let's take a look.
So we want rain to be true. We want traffic to be false. So rain is true.
Traffic is false. Yeah, that looks good to me. So 5 and 6, great.
And now what do we do? We want the probability that umbrella is false. So 5 or 6?
We just read it off the table. So umbrella is false here, so that means our red box is this probability.
So red box divided by blue box-- someone want to calculate that? Yep.
STUDENT: 0.2. PROFESSOR: 0.2. So it's 0.2 again.
Do we know anything else? What's next to that? STUDENT: 0.8.
PROFESSOR: 0.8, great. So this is pretty straightforward now. So we can fill in this table as well.
And if we do that, and the clicker works, we get these numbers. And we can check in our table, and those things will match.
So how do we check? Well, we could check by doing the same thing as before. Let's say we looked at this thing we want to check.
We look at false, false, false, given, false, false. So you're welcome to multiply those things out, but it should come out to 0.504, and this is really nice.
What else do we notice about this conditional probability table, probability of umbrella given rain and traffic? Any structure in there?
Yeah. STUDENT: It doesn't depend on traffic. PROFESSOR: It doesn't depend on traffic.
Yeah, so what do we notice? We see it doesn't depend on traffic if we know rain. So if we know rain is false, it doesn't matter whether traffic
is false or true. These two rows are the same. And similarly, if we know that rain is true,
it doesn't matter what traffic is. These two rows are the same. So essentially, this is telling us
we have conditional independence. So conditioned on rain, we know umbrella whether we know traffic or not.
So we can simplify this table. We can collapse this into this. So we take this, we collapse the two rows that are the same,
and we get rid of the traffic variable because it's the same no matter what. And then we condense that into a smaller table.
So this is the idea of conditional independence. Does that make sense to everybody? Questions so far?
Yeah. STUDENT: I have one question, the way you got 0.2. What-- did you just divide 0.014 divide by 0.7?
Is that-- PROFESSOR: When we got-- which 0.2? This one here?
STUDENT: So the denominator was 0.7, which is all the first four rows, but on the top-- PROFESSOR: Well, hold on.
So the first four rows, that might have been over-- wait. Hold on.
Which one are we looking at? So rain given traffic. So traffic-- so we have these two, and we have these two.
Sorry-- rain being false. So we were conditioning on rain being false, so we want the first four rows.
Yeah, so this should be-- what was this, 0.7? STUDENT: Yeah.
So why did you just divide 0.014 divided by 0.7? There's two rows that-- PROFESSOR: It should have been 0.14, not 0.014.
STUDENT: So it's sum of the-- PROFESSOR: Yeah, it's the sum of those two rows. Does that make sense?
Great. Other questions? Nice.
So we're going to collapse this big conditional probability table. It's not that big.
It's three variables, but now it's two variables. So we get some savings. We're going to keep doing this, and this is this idea
of conditional independence. So it's saying that we're going to exploit structure. We know that in the real world, like this idea
of strict independence isn't going to hold all the time, but maybe we can leverage the fact that some variables depend on other variables.
And then they don't depend on third variables, and we can figure out, if that structure exists, we can use it to our advantage.
So conditional independence just adds a third variable. It says X is conditionally independent of Y given Z if when we know Z, we don't need to know Y. So this probability
of X given Y and Z is essentially the same thing as probability of X given Z alone. And we can write it the other way as well.
We can say the probability of the joint distribution conditioned on Z is the product of the marginal distributions conditioned on Z So probability of XY given Z
is the probability of X given Z times probability of Y given Z. Another example, imagine that I have a toothache, and I'm trying to figure out if I have a cavity.
So this robot's going to help me to determine whether I have a cavity. It's going to inspect my teeth.
So we can model this with a three variable model. I have a toothache variable, a cavity variable, and a variable for whether this thing catches the cavity.
So if I have a cavity, the probability that the probe catches it doesn't depend on whether I have a toothache.
My toothache doesn't help this thing catch my cavities. And similarly, me not having a toothache doesn't affect this thing's ability to catch my cavities.
It depends on whether I have a cavity or not. That's what's going to affect this catch variable. So catch is conditionally independent of toothache given
cavity, and similarly, toothache is conditionally independent of catch given cavity. So we can write these in lots of different ways,
but there's all these conditional independence relationships for this particular problem. What about this domain?
So now we have a fire in our house, and there's a smoke alarm. And the alarm might go off if there's smoke.
What do we think is going to be conditionally independent here? I'll give you this hint. Yeah, keep your hands up so I know.
Yeah. STUDENT: Is it alarm only depends on the smoke to the fire because smoke would be coming in and out?
PROFESSOR: The alarm only depends on the smoke. So how would we write that? Conditionally independent or fully independent?
So probability of alarm-- STUDENT: Given there's a smoke. PROFESSOR: Given smoke.
STUDENT: And then-- PROFESSOR: And this is equal to the probability of the alarm given smoke and fire.
Do folks agree with that? Nodding heads. Yeah, so this is saying if I know that there's smoke
right next to my alarm, I'm pretty sure it's going to go off regardless of what kind of smoke that is. It could be smoke from a fire.
It could be smoke from a cigarette. It could be smoke from wherever. It could be smoke from my smoke machine.
If it gets close to the alarm, it's probably going to trigger the alarm. It doesn't really matter if it--
once I know that I have smoke close to the alarm, there may or may not be a fire. It doesn't really help me figure out
whether the alarm is going to go off. So yeah, we can model that this way. Here's another problem.
We'll call this one Ghostbusters. Imagine that I have a ghost that's hidden somewhere in this grid.
Right now I'm showing you where it is, but let's say that we don't know that the ghost is here. And what you can do is you can click on these squares,
and when you click on the squares, you get a color. The color tells you how close you are to the ghost. So if you get a red, you're right on the ghost, probably.
The sensor is a little noisy. So if you get a red, you're most likely on the ghost. If you get an orange, you're most likely
one or two squares away. You get a yellow, you're a little further away. Green is even farther than that.
So we're going to click on these squares, and then once you're confident based on the squares that you've clicked on and what colors you
got that you know where the ghost is, you can click this button to bust the ghost, and then the ghost will go away.
Does that make sense as an idea? So we have this noisy sensor. Sometimes it outputs yellow when it really
should be outputting green, but overall, if we get closer to the ghost, it's going to go from green to red.
So we can do a little demo of a system that's going to do that for us. So we'll start with uniform probability.
This is our prior distribution. This is before we-- prior to doing anything in this problem, we think the ghost could be anywhere.
It's equal across all these different squares. And then we're going to make a guess. And let's just say we guess here,
and we get a green observation. So if we get a green observation, what does that do to the probabilities?
It says green means that we're far away from the ghost. So that means that-- let's see what the best color is to use here, maybe orange.
So that means that in this region, we don't see a lot of probability anymore. So the ghost is probably hanging out in this corner or maybe down
here, one of these corners, not so much in this region. So you can see the probabilities adjust once you get this green observation even
though the sensors are noisy. So we can keep going. And as we keep going, we're going to make another guess.
And let's say that we guess over here, and we get this yellow observation in the corner. So what that does is it collapses our probability
distribution over possible places that the ghost could be, and-- oops, try not to advance too fast here.
So now we're looking at this zone because we know we're far away from the green, a little closer to the yellow, and so there's
this sort of ridge of probability that we might see the ghost in. Keep going, and so we're going to make another guess.
And then going to make another guess, and we're going to gradually try and narrow in on where we think the ghost is.
So first, we're going to guess there. We get an orange. Orange says we're getting even closer.
So now it's-- we're basically pretty sure we're at these two places on the ridge that's right next to the orange, and so we're going to guess one of them.
And we're going to get a red, and now we're pretty sure that it's right there. We get 0.97.
And just to be completely sure before-- we don't want to bust the wrong ghost or hit the wrong square. So we're going to guess a couple of squares
around that red square just to be sure that this wasn't sensor noise. But once we're sure, we're going to go ahead
and we're going to bust the ghost. And we'll hit it, great. Success-- ghost busted.
Yeah, so this is like if we had a 3 by 3 version of this problem, we can write down all the variables on the slide.
And we're going to have a variable for the ghost location. That's one of these nine positions. We're going to have variables for the color of each square,
so these are our observation variables. And the ghost can take on one of nine values. The colors can take on one of four values.
And so the total probability distribution size, the number of possible outcomes here is, there's nine for each--
nine ghost locations times the combinations of all the different colors. So we have-- get the pen--
we have nine. Oh, that's impossible to read. Let's try this-- nine times four different colors times nine--
to the ninth number of positions. So 9 times 4 to the ninth. And the way that the physics work here
is that we have this uniform prior distribution to start over the ghost location. And then our sensor model says when
we're trying to predict the color of a given square, let's say the color of square 1, 1. The probability that that comes out to be yellow
or any particular color is dependent on the goal, but it's not dependent on any of the other color observations. So once we know the goal, we don't
need to worry about the colors of the other squares. So this-- if we multiply 9 times 4 to the ninth, we get 2.3 million entries.
So this is a really big probability table. That's not so big for computers today, but as these problems get bigger and bigger,
it's going to be impossible to write out the whole joint distribution. So can we use independence here to help us?
Are C 1,1 and C 1, 2 independent? That's these squares here. So are the colors in these squares
independent of each other? Shaking your head. Can I ask why you're shaking your head?
STUDENT: Because if the ghost is far away from both of them, then they'll both be low. PROFESSOR: Yeah, so if the ghost is far away from both of them,
then they'll both be sort of that greenish yellowish color. And if it's close to both of them, they'll both be sort of orange or red together.
So you're going to see some correlations between these variables. You're going to see that knowing one of them
is going to tell you information about the other one. So if I see the orange, I'm pretty sure that the yellow is more likely
than green might be for that bottom square because they're only one apart. So what about conditional independence?
So we have this sensor model that says that the color of a particular observation on a square depends only on the distance to the goal.
So if I know that distance, if I know what the goal is, can I get away with this? If I know the ghost is here, now are they independent?
Well, they are because the orange one is-- we know how far away that is from the ghost, and we know how far away the yellow one is from the ghost.
And it doesn't matter whether we sample the orange one from our noisy sensor or we sample a green one or a red one. It's of no consequence to predicting
that yellow is probably going to happen in that bottom square. So now it's the case that yellow-- the probability of yellow conditioned on goal being 2,
3 is the same as the probability of yellow conditioned on goal being 2, 3 and knowing the orange square. So now we do have conditional independence.
So this idea of we have this big model-- it's exponentially large-- we could write that down as a product
of all these conditional probability tables like we did with the chain rule on the rain and the traffic example. But we don't want to do that because these conditional
probability tables are going to get just as big as our joint distribution when we get to the last one. This one has so many different variables in it.
But when we have conditional independence, we're going to be able to write this much more succinctly. Each one of these is only going to depend
on one variable at a time. So that's really nice. So we go from exponential to linear.
So that's the idea behind Bayes Nets, and this lecture is about how do we do that in practice? So we're going to take these big joint distributions,
and we're going to try to succinctly extract the structure so that we can succinctly write them down as just a product of small conditional probability tables.
Today's code is a Ghostbusters reference for those of you who've seen those movies-- ectoplasm. So everyone had time, almost?
It's on the website. So if you don't have the-- if you can't get the QR code, you can find it there.
But yeah, so the big picture is this. It's a technique that uses this concept of graphical models, which is really just models that use graphs,
in order to encode the structure of the world and allow us to write down these concise, small conditional probability tables.
So we're going to assume that each variable only depends on a certain small number of other variables, that there's this sort of really sparse structure
that's defining how these things are hooked up. And that's going to allow us to represent these really large probability tables with a small number
of conditional probability tables. Then we're going to be able to ask our model questions. So we can run these inference algorithms
and ask for things like what's the location of the ghost when I see all of these other observations? Or what's the probability that it's
raining outside when I see someone who has an umbrella? So the way that's going to work is we're going to use graphs. We love graphs in this class.
So we have, here's two different scenarios. One is where we're modeling the weather. The weather, we're going to have basically
these nodes in this graph are going to correspond to the variables in our scenario. So if it's just the weather, we have one node.
It's just the weather node. And if it's this cavity example, we have a cavity node, one for the toothache variable,
and one for whether this thing catches my cavities. And then the arcs between these nodes, these edges, are telling us about dependency relationships.
Or in fact, actually, what they're saying is if there's no arc between two nodes, then they are conditionally independent given their parents.
So more on what they exactly encode later. For now, you can imagine that these arrows just imply causality.
So the cavity causes the toothache. The cavity causes whether or not this thing catches my cavities. And so you can think of it as causal arrows.
Once I know this thing, I can predict the next thing. So go back to the coin flip examples, if I have n independent coin flips,
what would my graphical model for this look like? Well, we have n variables. So we can write those down, and then we
need to hook them up somehow. But they're independent, so we actually don't have any arcs. So this is the whole thing.
They're strictly independent, and so we don't have any edges between them. So this is what the graphical model looks like.
It's very simple for coin flips. Now we can go back to this traffic example. Let's say that I write down two variables--
traffic and umbrella. And I want to figure out how does the edge go between these? Does it go this way?
Does traffic cause me to take my umbrella? That doesn't feel quite right. Does me having an umbrella cause there to be traffic?
Be kind of nice if I had that much power with my umbrella, but I don't think that's how it works, either. So it really doesn't make much sense
until we add this third variable that kind of causes both of them. And so when we write this down, it's obvious.
Like, rain causes traffic. Rain causes me to take my umbrella, and so we can write down the graphical model like this.
If we go back to the smoke alarm, we said that fire causes smoke. Smoke causes the alarm, so then we
could write the graphical model like this. Or we could do this one for Ghostbusters which says that the ghost location tells me
the distribution over the colors that I'm going to see at each square. So the variables are the ghost's location, the different colors,
the different-- sorry, the different random variables representing the color that I observe at each location. And what we would like to calculate
is the probability that there's a-- of the ghost being at each of the locations conditioned on seeing the observations at all the squares.
This is called a Naive Bayes model. It's called naive because, in general, if you make this assumption that everything sort of depends
on one single variable, that's not going to work for you. But in this case, because of the way the sensor is designed, it'll work great.
It's not actually that naive, but this is known as a Naive Bayes model. We have one discreet query variable.
That's the ghost location, and then we have all these evidence variables, which are the observations.
And we're trying to figure out what the ghost location is conditioned on all the evidence. The evidence is conditionally independent given the query.
So that's starting to get a little more complicated, but we can build these things up pretty large. So if I'm an insurance company, I
might want to know what's the potential payout that I would need to make? What's the cost, the expected cost of medical liabilities
and property for my different insurance for a particular person that I'm thinking about insuring? So I can ask them a bunch of questions.
These blue things are things that I'm allowed to ask them. I can ask them about their driving record. I can ask them about whether they have their car in a garage
or what their age is. I'm not allowed to ask them what their socioeconomic status is. In some states, you're allowed to ask gender.
So you can actually discriminate on-- you can discriminate against people with their gender for some reason.
I'm not really sure why we're allowed to do that, but it is legal in some states. But other things are not legal.
So-- or you can't ask someone to-- if you ask someone are you risk averse? And you know that they're trying to get insurance from you--
you're the insurance company. You ask them, are you risk averse? What would you answer?
You're going to say that you're risk averse because that's going to give you a lower insurance premium. But you can ask someone whether they garaged their car
and whether they have an anti-theft device and what's their driving record. And you can kind of get at this quantity,
especially when you know their age and what kind of car they're driving. You can kind figure out what their risk
aversion is because maybe they buy a car that's especially safe. So you can get a sense of what this variable is based
on these observations, and you don't have to ask them the thing that you know they're going to lie to you about anyway. So all these orange things, you can't ask them about,
and all these purple things are the things you would like to try to calculate. And this is like-- this is what insurance companies.
Do they have models like this. You can also use it to troubleshoot your car if it won't start.
There's all these potential reasons why it might not start. Let's say you observe it's not starting. You can measure whether the lights turn on,
and you can try and get a sense of why your car's not starting. I'm not going to belabor that example. So the way this works, how do you build a Bayes Net?
We explained this a little bit, but let's get into the technical details. So each of these nodes is going to correspond
to one of the random variables, and we're going to arrange them in a directed acyclic graph. It could look like this.
And then we're going to have a conditional distribution for each node given its parents in the graph. So what does that look like?
We have this table. It doesn't have to be a table, but if it's discrete, it ends up looking like a table.
If it was continuous, then maybe it was-- maybe it's the case that the ghost determines the hue of this color, but maybe the hue
is like normally distributed in some range. So you could have-- this table could be like a mean and a variance of that Gaussian distribution
or something. But yeah, keeping with discrete for now, we'll have a table for each of these entries.
So G is going to have an entry for each location. It's going to be a uniform prior, and then each of these P of C conditioned on Gs
is going to have an entry for each location and an entry for each color. And then that's what our conditional table
is going to look like. So this is exactly the same thing that we saw before when we had the traffic and the rain
example, but now we have this diagram that's telling us which structure do we need to include so that we don't have all that redundancy.
So we're going to have one of those for each color. So a Bayes Net is this topology, the graph, and it's also this local conditional probability
table that's telling us how do these things relate to each other, and what depends on what. Let's try that with an example.
This is the idea of I have an alarm in my house, and this alarm is pretty good at detecting burglars. But it also sometimes tells me when
I have earthquakes in my house. But I'm actually never home, so I need my neighbors to call me when my alarm goes off
to let me know that my alarm is going off. So maybe I'm always working or something. So I have all these variables.
There's a burglary variable. Did it happen? Was there an earthquake?
Did the alarm go off, and then did I get a call from my neighbor, John, or did I get a call from my neighbor, Mary?
This example comes from Judea Pearl. So I can model that like this. The burglary causes the alarm to go off.
The earthquake causes the alarm to go off, and my alarm going off causes my neighbors to call me. So that's how the arrows go.
Does that make sense? Questions on this so far? Yes.
STUDENT: Is there any reason why there is two different neighbors? PROFESSOR: Is there a reason why there's two different neighbors?
Yes, there is, but we'll get to that in a second. So they'll have different probabilities of calling me. So let's write down some conditional probability tables.
This is just a structure. This isn't the whole Bayes Net. The Bayes Net is this plus we have
to go to these conditional probability tables. So the probability that there's a burglary maybe is like 1 in 1,000.
Let's say that this is where Judea Pearl lives in LA. Then there's the probability of an earthquake, which maybe is 2 in 1,000 for LA.
And then given these two variables, we can write down the probability that the alarm goes off.
The alarm is very likely to go off if there's both an earthquake and there's a burglar. It's also quite likely to go off if there's just
a burglar and no earthquake. If there's only an earthquake and no burglar, it's maybe 3 out of 10 it's going to go off.
And then if there's nothing, if there's no burglar, no earthquake, let's say that there's a 1 in 1,000 chance that a mouse bumps into the alarm and causes it to go off.
And then there's two neighbors. There's John and there's Mary. Let's say that John is 0.9 likely
to call when the alarm goes off and also 0.05 likely to call just for no reason, maybe listening to the TV, or there's some reason he thought the alarm went off.
Maybe Mary is hard of hearing, so she can't hear the alarm. And so there's only a 0.7 chance that she calls when the alarm goes off, but she's
less likely to call when it's not going off because she's better at focusing or something, I don't know. So we can write all that stuff down,
and that allows us to express this large conditional probability table using only these-- a small number of parameters.
So how many parameters? Well, there's one for each of these because, remember, once we know one of those values,
the other one has to sum to one. So we have one for each of these top two tables. We have two for each of these bottom two tables--
two parameters-- and then we have four parameters in this middle table, a total of 10 parameters. So if we were going to do this the hard way where we write down
the whole joint distribution, we have five different binary variables. So that's 2 to the fifth parameters minus 1
because they have to sum to 1-- so 31 parameters instead of 10. So already, for just five variables and just binary, this is saving us quite a bit.
So the Bayes Net semantics are essentially that this allows us to write down these conditional probabilities conditioned only
on the parents of each node, and then we can multiply those things together to construct our joint distribution
without any loss of fidelity. This exploits the sort of sparse structure. So we're assuming that the number of parents
is usually small. So I already sort of alluded to this, but the size of a Bayes Net is going
to be a lot smaller than the size of the joint distribution. So the joint distribution over n variables, each of which has d values, does anyone know what that would be?
Yes. STUDENT: D to the n. PROFESSOR: D to the n, great.
And what about an N-node Bayes Net if the nodes have at most k parents? And let's say they still have d values.
So I need a conditional probability table for each variable. They have k parents, and there's d
values that each variable can take on. What would one of those conditional probability tables be like?
Yeah. STUDENT: D to the k. PROFESSOR: D to the k, and then we have N of them.
So we need N times d to the k. So this gives you the power to calculate the whole joint distribution, but if you
have a lot of sparsity, if k is small, and you know what d is, then this goes from exponential to linear, which is really great.
It's also easier to write down these conditional probability tables than it is to write down the full joint distribution. Like if I said, what's the probability of it being rainy
and me having my umbrella and there being no traffic? Does anyone know what that probability would be? I think it's safe to say that nobody knows
what that probability would be. But if I say I know that I am more likely to take my umbrella when it's raining, and I know that it's
more likely for there to be traffic when it's raining, so now I can sort of define these conditional probability tables more easily than I could if I was trying to define
the whole joint distribution. So it makes things easier. And we'll see next time that it can
give faster answers to our questions about probabilities. So here's a question about a probability. What is the probability that there's a burglary
and there's no earthquake and the alarm goes off and neither John nor Mary calls? Stare at this for a second and see
what you think because we have all the answers here. We can construct the joint distribution from this information.
I'm not sure why there's a 33 on the slide. Ignore the 33. So what we can do is we can highlight
all the relevant quantities in these tables. So we highlight all these things. So first we know that there's a burglary.
Then we know that there's no earthquake. We know that the alarm does go off, true. We know that John doesn't call.
We know that Mary doesn't call. And so the probability of the burglary is 1 in 1,000. The probability there's no earthquake is 0.998.
Then this probability says condition on there being a burglary and no earthquake, what's the probability that the alarm goes off?
It's 0.94. Then we can get the probability that John doesn't call given that the alarm--
seems like that should be a true. I'm not sure why it's on the false. Oh yeah, because the alarm is true,
and then John doesn't call. So condition on the alarm being true, neither of those two people calls,
and so we multiply all these orange quantities together, and we get 0.000028. Nice.
So then we would just have to do that again for each of the 32 other values in our table, and then we would get our whole joint distribution.
So we can compare this idea to this idea of the chain rule, and we see that it looks very similar. Before we had the chain rule, we were just
multiplying by all the variables that came before that variable. Now we're saying, yeah, we just need to multiply by all of the parents of that variable
in this Bayes Net structure. So we can assume that these XIs are sorted into some topological ordering,
but the Bayes Net asserts the conditional independence of X from all the other factors that are not the parents of X. So we're saying X is conditionally independent of all
of its non-descendants. So once I know X's parents, once I know these two things, if those are the parents of X,
then now I know that-- let me use a better color. None of these variables influence X nothing up here
can influence X. X is conditionally independent of all of its non-descendants given these parents.
Yes, question. STUDENT: So would X be dependent on its parents [INDISCERNIBLE]?? Or is it dependent on [INDISCERNIBLE]??
PROFESSOR: It's conditionally independent of all of its non-descendants given its parents. So if you U1, and you know Um, and you know all the ones that
are in between-- all of these parents of X, like we could have-- there could be a U2 in here.
And that could also influence it. But if all of these things in this red and now blue circle, then anything in these other regions
that are outside of this sort of shield, this blanket, you're independent from conditioned on those parents. Yeah.
STUDENT: So if the one had a parent that was like [INDISCERNIBLE] that could be independent of U3 or-- PROFESSOR: X had another parent that's called U3?
STUDENT: U had U1 times-- PROFESSOR: Oh, U1 had a parent that's like-- let's call it V. So V is a parent of U. That's what you want to know?
STUDENT: Is that particular line descended? PROFESSOR: Yeah, it's not descended from X because it is an ancestor of X. Or it
could be sort of unrelated to X. So in the coin flip example, all the coins are independent of each other just because there's no connection between them.
Other questions? Nice. So let's look at this burglary example again.
We'll get rid of the John and Mary calling just to make the problem a little simpler. Let's try and figure out how would we
figure out the structure of this Bayes Net? Well, we could start by just-- we have the first variable, burglary.
There's no other variables yet, so we're sort of good. But now we want to add in the next variable, and we want to decide is this the relationship?
Does the burglary cause earthquakes? That doesn't feel right. So maybe we want to--
maybe say like the burglary and the earthquake are independent of each other. If they're independent of each other,
then we can say that we're good so far. Like we have drawn all the arrows that we need to draw. It also doesn't feel like earthquakes cause burglaries.
Although, I guess there's certain models where you might think that they would. We could add the alarm to the picture,
and now we say, well, now that the alarm is in there it makes sense. Like burglary causes alarms, probably.
Earthquake causes alarms, probably. And then we can write down these conditional probability tables, and then we're back where we were.
So what if we write them down in a different order? What if we start with the alarm variable? We can just pick which order we want to put them in.
So we put the alarm variable, that's fine by itself. But now we want to add the other variables, so we put the burglary variable down.
Is burglary conditionally independent of the alarm variable? Well, are they independent, just strictly independent?
No. I see people shaking their heads. There should be some kind of an arc here.
We don't know if it points this way or if it points the other way. And then if we add the earthquake variable in,
what's the relationship between the alarm and the earthquake? Are they independent? No.
So the alarm and the earthquake, again, we don't know necessarily which direction these arrows go in. But we know that there's got to be an arc between these two
things. And then what about burglaries and earthquakes? Well, we might be tempted to say,
we know that before, we had these arrows pointing in other directions, and like we knew that everything was fine. Why do we have to think about this arrow?
But let's say that I have this setup. If I find out that there's a burglary, does that tell me whether there was an earthquake
or not if I know that the alarm went off? So I know the alarm went off, and then I find out that there was a burglary.
Do I think it's more or less likely that there was an earthquake now? Well, probably if the alarm went off,
and I know there was a burglary, it probably went off because there was a burglary. So in some sense, this burglary variable
like explains away this concept of the earthquake going off. So if I have burglary, and I know the alarm went off, this like decreases my likelihood
of there being an earthquake just because the burglary is a sufficient explanation for the alarm going off. So these things are not necessarily conditionally
independent once I know the alarm because the burglary can explain away some of the earthquake probability. Does that make sense?
This is kind of a difficult concept to get sometimes. Yes. STUDENT: Would it apply in the opposite direction
to two different arrows? PROFESSOR: Yeah, so that's a good question. So you said it applies in the opposite direction,
so do you need two different arrows? So in a Bayes Net, you can kind of choose. The directions of the arrows, we've
said so far that they sort of represent causal directions like this one causes that thing. But it's really not quite the case
that that's what they signify. So you don't need to pick the right direction in the causal structure.
So it can be either direction. You could start with earthquake and then-- and do an arrow to burglary as well.
Really, what that's saying is that conditioned on alarm, burglary is not independent of earthquake, and so the arrow can go in either direction.
Those would be two different models, though. So we'll get to that in a second. So depending on which direction the arrow goes,
you'll get two different models. But so-- yeah, so we could write down our conditional probability tables here.
But what we're going to do when we try and write these down is we're going to run into some trouble here. So like the fact that we have these tables, like what's
the probability given that the alarm goes off of there being a burglary? Well, that kind of depends on the earthquake, right?
Like if there was an earthquake, then that changes the probability, so it's harder to write down this number. And similarly, if we know that the alarm went off
and there was a burglary, what's the probability of the earthquake? That's also kind of hard to think about.
So because we've written things down in a non causal direction, it's hard to estimate these conditional probability tables. It would be much easier if we had the right causal structure.
So here's another example of that. Let's go back to this traffic situation. Let's say that we only care about rain and traffic.
I'll get rid of the umbrellas. So is it the case that rain causes traffic, or is it the case that traffic causes rain?
So if rain causes traffic, then when the rain cloud arrives, it causes the cars to slow down because they don't want to hit each other.
If the traffic causes rain, then the minute this person stops being so close to the other car-- I don't know, the rain cloud gets
caught behind the wheel or something like that-- so this is a potential model of the world. It doesn't necessarily mean it's a good one,
but traffic could cause rain. We have this joint distribution, which tells us like we see rain and traffic together
3/16 of the time. We see rain and no traffic 1/16 of the time, et cetera. If we wanted to model these two different Bayes Nets,
we could do that. We could write down conditional probability tables. We could say that, well, if we just
say that rain is independent, it's its own variable, it maybe rains a quarter of the time. And then we have our conditional probability table
that says if it's raining, then there's 3/4 likelihood to exist in traffic. And if it's not raining, it's 50/50.
That sort of makes sense. But there's an equivalent interpretation, which is that traffic causes rain, and this set of tables
is actually totally consistent with the same joint distribution. So we can say, there's 9/16 probability of traffic
and 7/16 probability of no traffic. That happens sort of just independent of anything. And then once we know traffic, we know that if we see traffic,
rain is 1/3 likely. If we don't see traffic, rain is 1/7 likely. And if this is the case, you end up
with the exact same joint distribution. But it's highly non intuitive why these would be the numbers that you would see.
So these are two different ways that you can model the causality relationships between these variables.
Does that make sense? The arrows can go either way. We get to choose, so we've got to try and choose them
correctly. Questions? So yeah, this question of causality,
like the Bayes Net, when it reflects the true causal patterns, it's easier to write down the conditional probability tables.
It makes more intuitive sense. It simplifies the connections. We end up with fewer edges, often,
and it's often easier to elicit these probabilities from experts because experts kind of have a sense of how things work. And that model is a causal model, generally.
So they don't need to be-- the Bayes Nets don't need to be causal. We can write them down in either direction of causality.
But it's sometimes is harder to write them down, so like the traffic and rain example, or we end up with arrows that just reflect correlation and not
causation. Sometimes it's the case that there is no causal model of the domain because we're
missing some variables. So the idea here is that we have these-- we want to build up these Bayes Nets.
We want to encode all the structure of the world. That's going to help us with modeling these probabilistic questions.
But if we build them up in such a way that we don't capture one of these variables, that might be okay.
It might still be a good approximation of real life. But if we're missing certain variables, we might not capture all the causality relationships.
So what the arrows really mean, you can think of them as causality, but you should be careful because it might be the case that they
don't encode causality. Really, they encode conditional independence. So when there's an edge, you know that there's
a dependence relationship. And when there's no edge between two variables, they're conditionally independent given their parents.
So let me summarize so far. So we have this idea of independence and conditional independence.
These are important forms of probabilistic knowledge. They're going to help us to write down and express these complicated joint distributions
over large probability models. A Bayes Net is just going to allow us to encode that really succinctly so that we can
go from this exponentially sized thing down to a linearly sized thing. And when we do that, that means that we can store these things
in memory, and it means that we can query them more efficiently as well. So that's the next thing that we'll talk about.
But the most important concept that I think I want to try to communicate to you is that this allows us the flexibility
to trade off between model accuracy and this memory compute efficiency. So we started with this coin flip strict independence style
model. That's not a very good model. If we do that, we saw that even in the train--
in the traffic and rain and umbrella example. We could no longer model the joint distribution by assuming all the variables were strictly independent.
From there, you can add a little bit of structure. You can get to a naive Bayes model. This allows you to ask questions about some query, Variable A,
conditioned on a bunch of evidence. And then you can add more and more structure. You get these sparse Bayes Nets like the earthquake example,
and then you can get the full joint distribution where A affects everything, B affects C, D, and E, et cetera. And if you have this structure, this is the full generality.
But now you have this exponentially large probability table, whereas here, you just have a small number of linear conditional probability tables.
So this spectrum, going from strict independence to the full joint distribution, is kind of like the space of probability models.
And where you end up landing in this space is sort of up to you as the modeling person. Like you get to decide as the designer what is the right way
to model this problem? Do I have enough memory that I can afford the full joint distribution,
or do I need to restrict my model? And if I'm restricting it, how should I think about the probability and the causality relationships
and the conditional independence between the different variables so that I can model this effectively? So when we put our designer hats on and we're modeling the world,
this is a tool in the toolbox Questions so far? Yes. STUDENT: Can the Bayes Net be a cyclic one?
PROFESSOR: Can the Bayes Net be a cyclic one? No. We're going to restrict it to just be a DAG-- a Directed
Acyclic Graph. And the reason for that is you need to think about this idea of the chain rule.
So we have this joint distribution. What are we trying to do with our Bayes Net? We're trying to model this joint distribution.
So if I have A, B, C, D, and E, the joint distribution says that the probability of A times the probability of B conditioned on A times the probability of C conditioned
on A and B times the probability of D conditioned on A, B, and C times the probability of E conditioned on everything else gives me the joint distribution again.
So there's no reason when we have this chain rule that we would need there to be a cycle because as long as you go in order--
A, B, C, D, E-- then every variable can at most depend on all of its predecessors. And you can work out the chain rule,
and then you get this acyclic graph. So this is an acyclic graph that's got directed edges that captures the whole joint distribution,
and that's sufficient. So worst case, you can always write this down, and that's a valid Bayes Net.
Question. STUDENT: What's the difference between the naive and the sparse space again?
PROFESSOR: Great, so what's the difference between naive Bayes and sparse Bayes Net? So the idea with--
this isn't like an official name. This is just like the concept of a Bayes Net with sparsity. So the idea is, most of the time when we're writing down
Bayes Nets, we're hoping that each parent has a small number-- or sorry, each variable has a small number of parents. So that's that sparsity.
And then in the naive Bayes model, we're saying that there is exactly one variable that is sort of independent
of all the other variables or causally doesn't depend on any of these other variables. And then conditioned on that first variable,
which is generally the query variable-- you can think of this as where's the ghost in the Ghostbusters example--
we have all these other evidence variables. So this is like what did I observe at square 1, 1? What did I observe at square 1, 2, et cetera?
And the naive Bayes model is just set up so that there's one parent, and all these other variables are evidence variables.
So this is a type of sparse Bayes Net because all the nodes have, at most, one parent. But it's sort of making an even further assumption, which
is that there's one query variable, and the rest are evidence variables. So if we go back to the car insurance example,
this one, so we have three query variables. We want to say, what's the medical cost, what's the liability cost, and what's the property cost of me insuring
this particular person? And we have a bunch of evidence variables. We can ask them their age.
We can ask them how long they've been driving for. We can ask them what their driving record is. Do they have any accidents?
This anti-theft thing is nice because it tells us like, for one thing, we know that it deters theft. So if we ask them if they have an anti-theft system,
that tells us about this variable, which eventually will flow into the property cost. But it also tells us about their risk aversion.
So if they're the kind of person who would buy an anti-theft device, then we can go backwards and get this risk aversion variable
in some capacity. So this would be somewhat sparse Bayes Net. And then you can get more or less sparse
depending on how connected you make these different pieces. Does that answer your question? STUDENT: Yes, thank you.
PROFESSOR: Other questions? So we have about 10 minutes left, so let me-- what do folks want to do?
Do you want to stop there, or do you want to go on to inference, like how do we ask these models questions? We have 10 minutes left.
It would be helpful to keep going, or should we stop? Hands up for stop. Hands up for keep going.
cool. We'll do inference next time.
A Bayes Net is a probabilistic graphical model that uses a directed acyclic graph (DAG) to represent a set of random variables and their conditional dependencies. Its primary purpose is to efficiently model uncertainty by encoding conditional independence relationships, which reduces the complexity of joint probability distributions from exponential to linear scales. This makes it invaluable for reasoning under uncertainty in fields like AI, diagnostics, and risk assessment.
Conditional independence is a key property where two variables are independent given a third variable. For example, in the traffic/rain/umbrella example, knowing whether it is raining makes the traffic condition irrelevant for predicting if someone carries an umbrella. This reduces computational complexity because instead of storing probabilities for all possible combinations, you can break the problem into smaller, localized conditional probability tables (CPTs).
To construct a Bayes Net, follow these steps: 1) Identify random variables relevant to the system. 2) Draw directed edges from cause to effect to create a DAG, respecting causal relationships for easier probability assignment. 3) For each node, specify a conditional probability table (CPT) quantifying the node’s probability given its parents. The result is a compact representation of the joint distribution, where the total parameters are the sum of CPT sizes (e.g., 10 for 5 binary variables) versus 31 for a full joint distribution.
The 'explaining away' effect occurs when multiple potential causes can influence the same effect. Once one cause is confirmed, the probability of alternative causes decreases. In the Burglary/Earthquake example, if the alarm goes off, the probability of an earthquake is high, but if you also learn there was a burglary, the probability of an earthquake drops because the burglary 'explains away' the alarm. This is a form of reasoning that Bayes Nets handle naturally.
Bayes Nets offer a balanced approach: they are more realistic than strict independence (which is rare in real life) but far more computationally efficient than a full joint distribution. For instance, in the Ghostbusters example, a full joint distribution requires 2.3 million possible outcomes, while a Naive Bayes model (a simple Bayes Net) needs only 36 conditional probability tables. This trade-off allows solving practical problems like spam detection or insurance risk assessment without overwhelming computational resources.
A powerful application is car troubleshooting. The network includes nodes for causes (e.g., battery failure, alternator issues, starter problems) and symptoms (e.g., car won't start, lights dim). Edges connect likely causes to their symptoms. When symptoms are observed, the Bayes Net updates probabilities for each possible cause using Bayes’ rule. This allows mechanics to prioritize the most probable root issues, combining multiple evidence pieces efficiently.
While edges in a Bayes Net often reflect causal direction (e.g., Rain → Traffic), the graph strictly encodes conditional independence assumptions, not direct causality. The arrows can be reversed mathematically and still represent the same joint probability distribution, but using causal directions simplifies manual probability elicitation from experts. So, the net models 'what depends on what' rather than forcing a causal interpretation.
Keep this summary
Save it to LunaNotes and it becomes a real note in your library — editable, searchable, and ready to turn into flashcards or a diagram. Free to start.
Save to LunaNotesOr summarise for another video.
This summary and transcript were automatically generated using AI with the Free YouTube Transcript Summary Tool by LunaNotes.
Related summaries
Introduction to Probability: Foundations for AI & Bayes' Nets
Learn the core concepts of probability theory essential for AI, including probability models, random variables, joint & conditional distributions, and Bayes' rule. This lecture explains why probabilistic reasoning is crucial for handling real-world uncertainty in machine learning and decision-making, with intuitive examples and games.
Calculating Conditional Probabilities Using Tree Diagrams and Dice Rolls
This summary explains how to calculate probabilities and conditional probabilities using tree diagrams, including examples with bus arrival times, ball selections without replacement, and dice rolls. Key concepts such as intersections, complements, and conditional probability formulas are demonstrated with step-by-step calculations.
Probability & Statistics: The Ultimate Guide to Modeling Uncertainty (Course Intro)
Professor Steve Brunton launches an exciting new short course on probability and statistics. This introductory overview explains why probability is a foundational tool for data science and machine learning, provides real-world examples from thermodynamics to weather, and outlines what you will learn in the first 10 hours dedicated to probability.
Introduction to Probability and Statistics: Key Concepts and Terminology
In this video, Dr. Gajendra Purohit introduces the fundamentals of probability and statistics, covering essential terminology, types of events, and key concepts such as random experiments, sample space, and probability calculations. The session aims to provide a solid foundation for students preparing for advanced mathematics exams.
Understanding Game Theory: Key Concepts and Real-World Applications
In this lecture, Professor Ben Polak explores the fundamentals of game theory, focusing on the Prisoners' Dilemma and its implications in real-world scenarios. He emphasizes the importance of understanding payoffs, strategies, and the concept of rationality in predicting outcomes in competitive situations.
Most viewed summaries
A Comprehensive Guide to Using Stable Diffusion Forge UI
Explore the Stable Diffusion Forge UI, customizable settings, models, and more to enhance your image generation experience.
Kolonyalismo at Imperyalismo: Ang Kasaysayan ng Pagsakop sa Pilipinas
Tuklasin ang kasaysayan ng kolonyalismo at imperyalismo sa Pilipinas sa pamamagitan ni Ferdinand Magellan.
Mastering Inpainting with Stable Diffusion: Fix Mistakes and Enhance Your Images
Learn to fix mistakes and enhance images with Stable Diffusion's inpainting features effectively.
Pamamaraan at Patakarang Kolonyal ng mga Espanyol sa Pilipinas
Tuklasin ang mga pamamaraan at patakaran ng mga Espanyol sa Pilipinas, at ang epekto nito sa mga Pilipino.
How to Install and Configure Forge: A New Stable Diffusion Web UI
Learn to install and configure the new Forge web UI for Stable Diffusion, with tips on models and settings.
Found this summary useful?
Take it with you. One click puts it in your own LunaNotes library.
Save to LunaNotes