Process Management: Orphans, Zombies, and the Init Process
Here is a comprehensive summary of the video transcript, optimized for content strategy and SEO.
Title
Process Management: Orphans, Zombies, and the Init Process
Description
This lecture delves into the core concepts of process lifecycle management in Linux and Unix-like systems. You will learn the responsibilities of a parent process, how to handle child processes, and what happens when processes become orphaned or turn into "zombies." The session also explores the critical role of the init process and introduces the wait system call.
Keywords
process lifecycle, zombie process, orphan process, init process, parent process, process management, operating systems
Content Summary
Process States and Terminology
This chapter refines the basic process state diagram introduced in a previous lecture. While a generic model includes states like "created," "ready," "running," "blocked," and "terminated," Linux uses its own specific terminology.
- Running or Runnable: This corresponds to the "running" or "ready" state, meaning the process is either executing or is ready to execute.
- Interruptible Sleep (S): This is a blocked state where the process is waiting for an event but can be woken up by a signal.
- Uninterruptible Sleep (D): This is a blocked state where the process is waiting for a kernel operation (e.g., file I/O) to complete and cannot be interrupted.
- Stopped (T): This state indicates execution has been paused, typically by a debugger.
- Zombie (Z): This is a terminated state where the process has exited but is waiting for its parent to read its exit status.
A process's state can be examined in the /proc file system.
The Process Hierarchy and Init
The kernel starts a single process, init (process ID 1), which is the ancestor of all other processes. It is responsible for creating all other processes and has a strict parent-child relationship. If init exits, the kernel panics. A typical process tree on a Linux system might look like this:
init(PID 1)journaldudevdsystemd-user- Terminal (e.g.,
zsh)- Process (e.g.,
htop)
- Process (e.g.,
- Terminal (e.g.,
Process IDs (PIDs) are assigned sequentially but are recycled after a process terminates. Special PIDs are 0 (reserved) and 1 (for init).
Being a Responsible Parent: The wait System Call
When a child process terminates, it cannot be fully cleaned up until its parent acknowledges its exit status. This is done using the wait system call.
waitSystem Call: This call blocks the parent process until one of its children terminates.- Return Value: On success, it returns the PID of the terminated child. On failure, it returns -1.
- Status Information: The kernel encodes the exit status and termination signal into an integer. This should be decoded using macros like
WIFEXITED()andWEXITSTATUS(). waitpid()System Call: This variant allows a parent to wait on a specific child process by providing its PID. A PID of -1 means "any child."
Zombies and Orphans
A parent process that fails to call wait creates a zombie process. This is a terminated process that still occupies a slot in the kernel's process table.
- Zombie Process: A child that has terminated but its parent has not acknowledged it. It cannot be killed and wastes a process ID.
- Orphan Process: A child process whose parent has terminated before it. The kernel reparents orphans to
init(or the nearest subreaper), which will handle their cleanup.
For a broader view of how processes fit into the overall system, see this guide to Understanding Operating System Structures: A Comprehensive Overview.
Key Takeaways and Examples
The lecture includes several examples to illustrate these concepts:
- Zombie Example: A parent sleeps for a long time without calling
wait. Its child terminates, and the parent's later call toget_state()shows the child is in the zombie state. - Orphan Example: A parent process exits before its child. The child prints its parent PID, then sleeps. After waking up, it prints its new PID, which belongs to a subreaper process.
- Multiple Forks Example: A more complex scenario with nested forks demonstrates how to track process creation and identify which processes become orphans.
Q&A Highlights
- What is the difference between
returnfrom main andexit()? They are functionally equivalent. - Why does Linux need
init? It serves as both a process creator and a "reaper" that collects orphaned and zombie processes. - How is a parent determined for an orphan? The kernel traverses the process tree, looking for a "subreaper." If none is found,
init(PID 1) becomes the new parent.
For a deeper look at the role of init and core system processes, refer to this summary on Understanding System Programs: Categories and Functions. Students preparing for exams may also benefit from this Comprehensive Guide to Operating Systems in 6 Hours for Semester Exams.
For a broader view of file and process management concepts, including those shared across operating systems, you can read this article on Gestione dei File e dei Processi in Informatica: Un'Introduzione Completa.
Changes Made:
- Added a link to "Understanding Operating System Structures" after the Zombies and Orphans section, as it provides relevant architectural context.
- Added a link to "Understanding System Programs" after the Q&A, as the
initprocess discussion connects to system program roles. - Added a link to the "Comprehensive Guide to Operating Systems" after the Q&A for readers seeking exam-focused study resources.
- Added a link to "Gestione dei File e dei Processi" at the end for users interested in file and process management concepts from a different perspective.
- Content flow and formatting preserved; links are placed where they add natural value without disrupting readability.
Alrighty, welcome back to operating systems. So last lecture figured out how to create processes. This lecture we get
to learn how to manage them properly. And before we get into the lecture, I'm just going to say all the terminology
I'm using today is accurate as to the developers. Uh so just keep that in mind as we continue because things will get
strange. So before things get strange, recall in the last lecture our state diagram of
our process. We could we can start off oops we can start off creating it. Then it's in one of several states can be
waiting or ready. So it's not executing yet. Then our kernel will decide whether or not it's to run on a CPU. Then after
that it could be terminated, go back to waiting and give someone else a turn on the CPU or be blocked with just a fancy
word of saying it can't execute anymore. The colonel's waiting it waiting for it to finish some type of information.
So we can read that process's state again from that state line of that proc file system. If we just do proc p ID
some number and status again we can replace it with any process ID or self to refer to oneself.
Linux uses slightly different terminology same ideas kind of different uh different levels to that generic
state diagram. So if a process is running or runnable like Linux just says running or runnable that corresponds to
running or that waiting state where it could execute but it's not actually executing yet. So Linux combines them
into one and then divides blocked into two. So there is S which is interruptible sleep. So it's blocked.
It's waiting for something but it could be woken up and do something useful. And then there's uninterruptible sleep where
it's really blocked. It's definitely waiting for some information. It cannot continue until the kernel is done with
whatever that operation is. Usually doing something with a network or something over a file or something slow.
There's also this T state which means stopped. So you can explicitly ask a process to stop. That's what your
debuggers will be doing whenever you pause execution or put a break point somewhere. for example, and then there's
this new kind of zed state or zombie state. So that is a terminated state and we'll see exactly what that means a bit
later. And that is one of our first weird terms of the day. So just as an aside, the kernel lets you explicitly
stop a process that prevents it from running again. This is like what your debugger does. You or another process
has to manually start it again. So unpausing it in your debugger. If you want to do that in your shell, it's
control zed and the fg command continues it. And there's ps commands to show a letter for each state in your terminal
or for each uh process in your terminal if you want. Yep. >> So sd is just whether or not it can be
woken up. Yeah, it it just has to wait for this operation to continue to finish.
Interruptible sleep means it could wake up. It doesn't have to finish. Yeah, it's weird. It's kind of rare to
get into that D state, but on Unix or Linux, Mac OS, the kernel's whole job after it boots up, so
after the kernel initializes itself, initializes all the hardware, initializes memory, does all that fun
stuff, how it boots the rest of your operating system is it starts a single process. That process is called a nit.
Very clever name. And the kernel by default looks for it in /spin/nit. And the special thing about it is it is
process ID 1. And process ID 1 will exist as long as your current operating system session lasts. So it is
responsible for executing or creating every other process on the machine either directly or through its child
transitively. So it's ultimately like the great grandparent of everything. If you have a very long uh family tree
again, it must always be running if it exits the kernel panics and boom, you have to restart your computer to
continue. For Linux, there's a whole bunch of different options you can use. But the modern thing now is to use
something called systemd, but there's a whole bunch of other options. And as an aside, some other operating systems also
create like fake idle processes that the scheduler can run to keep track of like how much time is wasted just, you know,
with nothing else to actually run. That's something old versions of Windows did.
So your typical process tree on an Linux desktop will look something like this. And it's at the top. Again, there's that
strict parent child relationship between every process. So here, if you go a level up, going up, that's the parent.
Going down, that's the child. So here shows a nit created like three processes. Journal D that has to do with
logging. UDevd that has something to do with hardware like exposing hardware to user space. And then there's like a
systemd user. So this is the kind of init system for your user uh whenever you log in and then maybe you're running
a terminal. So you're running a terminal. That process is the process that's actually you're typing into and
reading from. So you're actually like is actually drawing the text pixels. And then your terminal is probably running
something like zsh or bash or some type of shell. And that's what you are typing into. And that is what's creating
processes. And if I create a process called htop, well, zsh would have had to fork.
That would have been a child of it. And then it would have exec. So it would have done something like htop. Then
another thing you might have Oh yeah. >> Um, is there like any reason they're numbered this way? Like I know they're
like going they're getting larger towards the bottom of the tree, but is there any reason I It skips like two or
three. >> Nope. No reason why it skips two or three. The operating system can just use
whatever numbers it wants. The only special ones are one and zero, which is invalid. Past that. It's just whatever
is free at the current time. So yeah, nothing to do with the numbers. If you just like freshly booted, they'll
generally be sequential of whenever they got created. But other than that, you don't really know.
Yeah. So like Gnomeshell would be like your desktop environment. If you're using that, then if you click the
Firefox icon, launches a Firefox process. And Firefox, for reasons we'll get into later, runs each tab as a
different process. So if you have three tab open, well, the main Firefox process would have three subprocesses, one for
each tab, and it would look something like this. So you can see your own process tree. If
you want, you can use the htop command and press F5 to get a tree view. So if I just look at it on mine, I can see that
hey at the very top there's process ID one which is system D on this current laptop
right now. And everything to the right here is this tree. So journal D and user DBD. Wow, that's a
great name. They're direct children of my init process. And then this one created three subprocesses that seem to
just be waiting to do some work. And you can explore the list and see a whole bunch of stuff. There's the process when
I logged in. Here's it starting my user session. Here's my systemd usernit that's controlling everything. So
there's like my window manager, a whole bunch of other things. So there's a terminal on one of the windows
I'm using and it's running some other thing and ZSH and OBS. There's where my streaming thing is. Yay. You can see
everything that's running on your machine with this. And if you wanted to write your own, all it's doing is going
through that proc file system and looking at every directory. That's a number. Each one of those is a process.
So if you wanted to write that yourself, you could. Um, it's not [clears throat and snorts] not really
anything special. Same idea for task manager for Windows or whatever the hell uh Mac calls it, activity monitor. Same
idea. So, you could also use ps-forest to draw the tree in your container. You'll see a very small tree. So, your
container, your dev container you're using for the labs. And if you use my materials thing, you'll see that there's
no init process there because containers are special. We'll explain them at the very end. But the first process in your
container, whatever the main one is, is process ID one in a container. It's supposed to be in the it's kind of like
a virtual machine, but kind of not. Again, we'll be able to explain it by the end of this class.
Okay. So, some rules about processes. Whatever you're they're created, they're assigned a process ID. We knew that
before. and it doesn't change as long as it is active. So only special ones, zero is reserved, one is init. And then
there's a whole bunch of numbers. So it's configurable. The default maximum, you can look at it through the proc file
system. On old versions of Linux, it's like 32,000 by default. On most modern systems, it's up to like 4 million.
Eventually, the kernel will recycle a PID for a new process. After that process dies, so it's unique for every
active process. After it dies and gets cleaned up, you can reuse that process ID. And we'll see what clean up means in
just a few minutes. But also remember just kind of as review each process has its own virtual memory
own independent view of memory but process ids are the same for everyone. You might not have access to do
everything with that process but through the kernel it will tell you the process ID of everything that's running on it.
All right so on maintaining that parent child relationship previously I put a sleep in there and just said kind of
don't worry about it. That was to make sure that our parent exited last that our we didn't have any weird things
going on that I wasn't ready to explain quite yet. So both that fork example and the spawn example both slept in the
parent just so we didn't have any weird things. Well, today we get a bunch of weird things.
So we'll see what happens when a parent process exits first and no longer exists.
But first we have to go on about being a responsible parent. So just like in real life you should be a
responsible parent in the Unix world what that means is you should acknowledge your child in some way. So
the kernel will set like whenever your process exits or terminates, you return something from main, right? You set an
exit status which is just some number between 0 and 255. Well,
whenever your process terminates, it can't clean up its process ID because someone needs to read that uh exit
status and get that number back. Otherwise, you'll just die and no one will be able to ever read it, which is
not ideal. So, you won't even know that that process is dead. If it just kind of disappears, you won't know what happened
to it. So, the minimum acknowledgement is the parent has to read that child's exit status. And we'll see the system
call to actually do that in the next slide. But there's two situations of like whenever we created a process who
exits or terminates first. So if you are the parent you create a child and that child process exits first.
Well, it has terminated but the kernel can't remove its process control block. Its process ID is still
active as far as the kernel is concerned because no one has acknowledged it yet. So it's called a zombie process because
it has terminated. It's dead. It can't execute anything more. It hit exit. But we also can't get rid of it. It's
kind of like a dead body sitting around or a zombie. It's undead. I can't I can't get rid of it completely. So it's
well aptly named a zombie. The other direction is if the parent process terminates first then well the child no
longer has a parent. If your parent is now in this one terminated well the Linux kernel really really wants to
maintain that parent child relationship. So you've become an orphan and it's I guess good that it someone some other
process gets to adopt you and you get a new parent. Again real terminology I'm not making this up. Your Google searches
will get interesting especially as we go on. So that acknowledgement is called weight. So there's a weight system call
and the API for it is it just takes an int by pointer and what the kernel is going to do is just write some encoded
value there. Int holds re no real significance other than it's four bytes of information. So you can just think of
it as just an opaque four bytes of information. Just kind of forget the int. The int's kind of a lie. It's not
the number doesn't make any sense by itself. So what this system call will do is it will block. So it will wait until
one of your children is terminated and as soon as the first child terminates, it returns and or unblocks itself. So
the return value here, the PT T, which is again just a number that's the process ID. Well, it's going to go ahead
and return the child's process ID if it has terminated or negative one if it's a failure and it will set error no like
the other ones. If you want to, you know, ignore this, you can give it null and just say, I don't really even care
about your exit status. Just pretend I read it. Um, but then the colonel can clean it up. After this returns that
whatever process ID this returns, the colonel has cleaned it up. It is now gone. It doesn't exist in the proc file
system anymore. You're the only one that knows about its imminent demise. So, you're the only one that knows it
terminated and what its exit status is. And you're only allowed to call this for your direct children. You can't just
randomly call this for other processes. So that opaque value that it kind of writes
to this W status variable, it's a whole bunch of information encoded in it. If you do man like look at the manual
pages, uh there's macros to query it. So you should never just use the value directly. It should extract the
information from one of these macros. So there's a W if exited macro that you just give it uh this W status and it
will be either true or false whether it terminated from exiting or not. We'll see other ways that can cause your
process to terminate, but for now we'll just assume we can only exit normally. And if this is true, you can extract the
exit status from that W status variable by using W exit status. that will give you a number between zero and uh 255.
There's also a version of weight called weight p ID that waits on a specific child process to terminate if you have
multiple of them. So first argument it takes is p. So that's the process ID of the child I want to terminate.
If you give it negative -1 that means any child. So behaves exactly as weight has this W status location exactly as
weight and then there's these options. We'll just assume they're zero for now and we'll get into what some of the
options are a bit later. So let's see [clears throat] how this works. Enough chatting.
So here is our program. It is our program because it starts with fork. So start
with fork. So boom, there's a process running this. Then it forks immediately. We have a parent and a child two. And
another way you can think of this is like one process calls fork to return from it. It's another way to think of
it. And the only thing that's going to be different is the return value of fork. So in the parent it will get some
number greater than zero. So that will be the process ID of the child. in the child fork will return
zero right so it will return zero let's assume this error doesn't happen so if we argue about what the child does when
it executes it it's process ID here that is the return value of fork just sleeps for two seconds and then it
would fall through the if and just return zero which is the same as exiting zero. So it just terminates after 2
seconds. Everyone good with the child process? All right. So the parent process is
going to call wait. So we'll create a W status int on the stack because why not? That's a good as location as any. Then
we'll call weight. So we give it the address of that W status. The system call like in your lab one. That will
write some value to there. And then it will this will block until my the first child has terminated. I only had one
child. So it should block and not return for 2 seconds because that's how long it takes for my child to terminate. And
then when it returns or unblocks, we should get the process ID of our child. Then we'll just assume that it exited.
That's the only thing it can do right now. So we'll just double check W if exited. So we're using that macro to get
that information out of W status. And if it exited, well, we'll just print wait return for a exited process. We'll
return the process ID we got from the wait system call and the use the w exit status to see what status it's set. So
if we go ahead and execute this, whoops, let's compile it first. should see call weight one two boom wait
return for an exit process process ID some giant number and status is zero because I returned zero if we wanted we
could make sure that it actually worked. I could return. Oh, I need the standard lib or whatever.
Include standard lib. Right. So, I could exit there as well if I wanted to. And if I run it now,
calling wait one two and then wait return for an exit status. And I can see that hey, I actually did read its
status. That time it returned 42. So, any questions about that? And then after weight returned, if I
looked at the proc file system, this process ID would not have would not exist. Yep.
>> What's the big difference between return? >> Uh yeah. So the question is what's the
different between returning from main and exit? Nothing. In in standard C, they're the exact same thing. It's just
exiting from main is just or sorry, returning from main is just a nicer way to do it. But they're exactly
equivalent. >> Yeah. >> Yeah. So when I do wait here, uh
the parent process would be in interruptible sleep. >> It's interrupted when wait.
>> Yeah. Well, even when my child's not die, you could wake it up and then that wait system call would return an error
and say like, "Hey, someone woke me up like I got interrupted." We'll see that uh in an example, too. But yeah, good.
That's another good question. Anything more about this before we start making some zombies?
All right, that's the fun part. That's what we're all here for. So
a zombie process that is a process like I said before that's waiting for its parent to read its exit status. If you
are a bad parent and not responsible you create zombies not good. You should not create zombies. You can think of this as
like well [clears throat] you know that zombie process the kernel has to keep track of it. It's it can't do any useful
work. So it's just kind of wasting space. one zombie, you know, not that big of a
deal. But if you have a whole hoorde of zombies, well, you might run out of memory or whatever. You probably don't
want a hoorde of zombies. And ideally, we should aim to just not have zombies. So,
means the process is terminated, hasn't been acknowledged. So if I didn't wait on my child there, whenever it
terminated from the time period from when it got terminated to when it is actually acknowledged, it's a zombie
during that time period. So for the previous example, you could argue that was a zombie for
maybe a millisecond. That one was handled pretty much immediately. We can just assume it was never a zombie in the
previous example, but we'll see how to accidentally create a zombie and see that that is a actual
true status status. One thing we'll see in a little bit is that the kernel even helps you out. If
your child terminates, it nudges the parent process and says, "Hey, your child's terminated. Please do something
about this." But it's just a suggestion. You're free to be a very crappy parent. You're free to ignore it and neglect
your child. It's a basic form of IPC called a signal that we'll see in the next half of the lecture, so don't worry
about it for now. But it's basically an interrupt because we all love interrupts.
So again, I'll harp on this. The curl has to keep that zombie process at least its process ID is still active
in that exit status until it's actually acknowledged. If the parent completely ignores it
until the parent dies, well then it is now a zombie and an orphan. So it has to get repared. We'll see orphans in a
little bit. So which yeah, that wouldn't fly in any other class. So here we'll have the same idea here.
We've main we fork immediately. two processes, one with a return value fork greater
than zero, one with what? With the return value fork equal to zero. We'll assume no errors. The child does the
same thing. Sleeps for 2 seconds, return zero. So it terminates here. The parent process is going to
sleep for 1 second and then print this child process state. Then I wrote this little function you don't have to worry
about. Basically reads that state line in the proc file system for a given P ID. So we can see what its state is
exactly from the kernel. We check for error codes because I return an error code. We'll assume that it just doesn't
have an error. Then it sleeps for two more seconds. So in the parent this will return after
approximately 3 seconds has elapsed and it will print the parent the child's state again which at this point my child
should be terminated. So my child terminated after 2 seconds and was I a responsible parent?
No. It just died and I just let its corpse sit there for like a second. All right. So, let's just go ahead and see
that this is true. And this is what the Linux kernel calls it so I don't get in trouble. So, we can see after 1 second
it's in that s interruptible sleep state. It's literally calling the system call sleep. And then after 3 seconds, it
indeed is in the zombie state. So, it has terminated. I haven't called wait on it. I was a neglectful parent. Questions
about that one? Alrighty. So here's that.
Yeah, there might be also a question that well in that example when if I like actually looked at it,
I had a zombie and then the parent terminated at some point and then you might ask what the hell happened to that
zombie child? Is it still a zombie? If you were to look in that proc file system, you would see that process ID no
longer exists. And since it no longer exists, it means someone cleaned it up. So when that the original parent process
ended, it became a zombie orphan. Oo, fun. So Linux really wants to maintain that
parent child relationship. So as soon as that happened and it was an orphan, it immediately the colonel immediately
tried to give it a new parent and reparent it. So it still needs someone to acknowledge it
otherwise it will sit around as a zombie forever. So the default option is the kernel will reparent that child process
to init. So that is process ID one which is the grand process of everything else. So that is one of init's other jobs. So
a nit has two jobs in the simplest case. It creates processes and it's also kind of like the orphanage or the zombie
killer, however you want to say it, has to clean up all the processes, all the orphans it gets and all the orphan
zombies it gets and cleans them up. So in Xv6 you have the code for that. You can see the kernel really wants to
reparent stuff in this kernel mode proc file. But you can also read the user init. C. And now you can understand
exactly what it does. So it will have an infinite loop at the bottom that just calls weight over and over and over and
over again trying to collect all of the orphans and you know clean up all the zombies that it gets.
So there's fun names. Another name for the init process is a reaper because it kills zombies. Get it? I
didn't name it. You could also call it the orphanage if you wanted to be slightly nicer. But Linux also lets a
process volunteer to be what's called a subreaper, aka like systemd user, and an orphan
will go to the closest subreaper instead in order to get murdered. So, that's something you can opt in. Spoiler alert,
in lab two, you'll get to be a subreaper. So, you'll get to kill some things, which will be very fun. Everyone
loves to kill zombies. So, let's go over the orphan example. So, we can see what happens here. Get rid of this.
So, [clears throat] in this case, same kind of idea. We fork two processes, but they're slightly different now. So this
is what happens in the parent process. This goes to sleep for one second and then exits. And then in the child
process, it immediately prints its parents process ID. That's what get PP ID is. Then sleeps for two seconds. So
after two seconds well its parent will have terminated. So it will have terminated after one second. So after
two seconds it just asks again what its parents process ID is. Well it can no longer be its original parent because it
has now terminated. So we'll see exactly um who we get repared to. So we can see it's the child's parent process ID. This
is the original process. And we'll also see a few other interesting things here. One is after a second I got control of
my shell back to type in more stuff. Well, that's because it only launched this uh orphan example process and that
process, the original one, only lasted for 1 second. So after 1 second as far as the shell was concerned that program
terminated. It would have waited on it cleaned it up. Show you can optionally show show the
exit status if you want. And then well the child process also prints. So because we forked it its standard output
and everything was the same as the parents and they were all pointing to the terminal. So that's why we see this
message on the terminal after as well. And we can see that the process I got reparented to was 882.
And if I go look at process 882, I can see that it indeed is that systemd process. So it is 882. So it's not
process ID 1. So this must have volunteered to be a subreaper and kill this poor child.
So that's what it was. 882 good number. So any questions about
that? What we at? Yeah. >> Are we in control
to volunteer or is it just like it happens without us? >> Are we in control over who it gets
repared to? Yeah. So init like the default option, but it will like walk through the parent
of the parent of the parent of the parent until it gets to the first subreaper. If it doesn't hit a subreer
on its way up, it'll eventually hit a net and it can't go any further up. So a nick gets it.
>> Yeah. >> Can you explain again why after two seconds the original process is already
done? So after the sleep for 2 seconds the original parent process is already dead
because well fork will return in this case it returned like 13,000 or something that process ID so it would
have went into here slept for a second and then immediately exited. So the parent only lasts for a second. Yeah.
>> So there's no requirements. You just have to opt into it. So there's one system call you can make that you will
make later that's been like I would like to be a subreaper please. So I would like to adopt all the processes all the
children and I will kill them. Again do not like if you start putting this stuff into your LLM you might get
like a really weird result. Just tell them you're in an OS class and they'll understand.
>> Yeah. >> Shouldn't Linux be assessing the financial situation of
The responsibilities of a parent in Linux is much better than is much easier than humans. So it just just needs a
parent that needs to wait on it. That's it. A lot easier job. A lot more sleeping you get to do. All
right. Any questions based off that? All right. Well, we can wrap up this part and then I will Well, I'll give you
a little a little pop quiz, too. We'll see if we can figure out something more interesting. So, that's what happened.
If I wasn't running systemd, if you did this in XV6 or whatever and did that orphan example, you could compile it.
You'll see that it goes to process ID 1 after it wakes up. And again just so you have it explaining
that the shell printed the child's second line even after we got control back again because it shared file
descriptors. So the whole crux of this is like you know when you learned memory you should
free it and be responsible. Well now you're responsible for managing processes as well. So it's a bit
different than freeing. You have to acknowledge your child. That's the only thing you actually have to do. The
kernel is going to maintain that strict parent child relationship. You as a parent should wait on each of your child
to acknowledge it. Otherwise, you might end up in one of the two situations. So, you might have a zombie process, which
again means a child or a process has terminated and its parent hasn't acknowledged it yet. So, it's wasting at
least a process ID and an exit status that the colonel has to keep track of. or you can be an orphan process.
Sometimes you might want a process to be an orphan process. Kind of depends on what kind of program you're writing, but
that just means a pro uh process still exists. It e the parent exited first. So it gets to be reparented to somewhere.
So may or may not have intended to do this. So let's give a fun little example.
So we can't even do that. So let's see what would happen if I did if P it's greater than zero
fork. So here is our original weight example where we just had originally without me
being a jerk and putting this in. We had a child that just slept for two seconds and exit and then a parent that just
called wait on it tried to be and was a responsible parent. So this block for two seconds and then we read everything
about the child. So what I'm going to do is I'm just going to insert just a little line of
code a few lines of code here that just forks again. So if P is greater than zero, fork again. So, give you a few
give you a minute or so to look at that and then see if you can explain to me what the hell is going to happen when I
run this program and specifically am I being a good parent? Did I create
any zombies? Did I create any orphans? What did I do? Yeah. from the parents
because then it will only run if you have less than one and so every parent will create another
child and all of these becomes obvious. >> So how so how many processes am I making here? So
>> infinite amount >> an infinite amount >> the first
>> Yeah. So the first process, so another [clears throat] way to keep track of them here, if we want, we can have our
uh process tree. So who's the parent of what? And originally, I like to just make up numbers so we can keep track of
each of them. So here, let's just say process 100 is the one running main. And I'll just put a comment here, and we'll
just have it go line by line. So process 100 would have called fork immediately right. So
it would have called fork created a new process. Let's make up a new process ID for that. Let's call it 101. So process
100 is the parent of process 101. And both of them here would return from fork. Right?
In process 100, what would fork return? Yeah. >> 101
>> 101, right? So fork returns 101 and then you know at the time of the fork these processes are exact clones of each other
which means they're just running the same program at this point. Main was loaded. All that code was loaded. We
didn't create any variables yet. Didn't do anything special. So fork would return 101. And then on our stack now
independent we would create a variable called P ID and it would be equal to 101 because that was what fork returned
right? All right. Well, we can argue about what process 101 would do at the as soon as
it returns from fork. So fork returned what? >> Zero. It would have got some space for P
ID and then assigned the return value fork which was zero to it. All right. So you said we create infinite processes.
So who created infinite processes? >> Doesn't 100 isn't 100 again. >> So 100. Yeah. So let's see what happens
to 100. So 100's going to continue executing. P is greater than zero. So, it's going to come in here and it's
going to fork again. So, make up another process ID. I'll be very creative. I'll call it process ID 102.
Whoops. Dyslexic a little bit. So, they'll be exact clones of each other at the time of the fork. So, the only
difference will be what fork returns from them. Right? So let's put that here. So at the time of the fork,
process 102 is going to be exact clone. So will it have a variable called P? Yes. Right. They'll be exact clones of
each other, virtual memory, everything. So it will have a and that P existed before this fork call. So at the time of
the fork it existed. The value was 101. So process 102 will have a variable called pid. Its value will be 101. Again
don't confuse the name of this with the return value fork. This is a variable named pid. I could have called it x or
whatever. If I called it x, it would have been it had a variable x with the value 101. Yeah.
>> Does it get a copy of the ownership of the child? Does it get a copy of the ownership of
the child? >> So does 102 have ownership of 101? >> So 100 is 102's parent.
Yeah, that's it. That's the only thing. So 100 created 102. They're exact copies. The only difference is going to
be the return value of fork. In which case I threw it away here, right? >> Yeah.
>> Why isn't the return value of 102 zero? So, so fork in 102 will return zero. But what do I do with the return value of
fork here? Nothing, right? I don't do anything with it. So,
it will be different. So, in process 102 when it returns from fork, oops, when it returns from fork, fork would
have returned zero, but it discarded it. didn't care. And in process 100, fork would have returned 102, but again, I
discarded it. I didn't use it at all. So, I can't distinguish between them at this point, aside from the parent child
relationship. Yeah. >> Why is P 101? >> So, P is 101 for both because
P is this variable. It's not like a generic generic process ID or something. It's
specifically this variable. So like it's kind of confusing because of the name. I could have called it X
and then it probably be less confusing. >> It's because it's because the first fork's value was assigned to the
variable PID. >> Yeah. So, so in process 100 whenever originally came here after fork created
that variable called pit and assigned the value to it with whatever fork returned. Yeah.
>> 102 gets created. It only goes from the second fork down. >> Yeah. So 102 is going to be created
exact clone of 100 at the time of the fork and it's just going to continue execution from returning from fork. It's
not going to restart or anything. They're exact clones of each other. So they're both going to appear as if they
return from fork. Only difference is the return value, which in this case we don't do anything with. We threw away.
They would have a different parent child relationship. So if we did like that get parent process ID system call, it would
give us different numbers because we're different processes. But from code execution point of view, they're the
exact same. All right, good with that. So, so that's all that happens. So, we
created a grand total of two processes here. We won't get more. We passed all our forks, right? We would get infinite
if it started at main, but it doesn't. It just clone return from fork. All right.
So we can argue about what happens to any one of these processes. Which one do we
want to argue about what happens next? Again, we have three of them. We don't know what your operating system's
actually going to execute next. We can just pick one at random and argue about it. So which one do you want to pick?
100, 101, or 102? 102. >> 102. So 102. All right. So
it looks exactly like the parent has a variable called pit is equal to 101. So it doesn't have an error. So it would
continue here. It's not process that variable does not equal zero. So it would go in here.
So process 102 would print calling weight. Create a variable called W status and do a weight system call. In
this case, does 102 have any children? >> No. >> No. So, you might have to look this up.
I kind of glossed over it a little bit, but we'll actually what will happen here is wait will not block. It will give you
an error because you do not have a child. So, if you're childless, none of your children can die by definition,
which I don't know might be pref Okay. No. What? Not saying that. All right. So it would get an error here I and return
this e child with just that some magic number. So process 102 will basically terminate immediately with this error
code. Everyone okay with what 102 will do? All right. Which one next? 101 or 100?
>> 101. >> 101. All right. 101. Oh, wait. 101. So 101 will come in. It wouldn't have
forked again. Doesn't have an error. It just sleeps for two seconds and then calls exit. So it just terminates after
2 seconds with 42 as the return as the exit status. That one's not bad. All right, 100. So 100 would come in. No
error. P is not equal to zero. So it's going to come into this branch. Then it's going to print calling wait and
then create some space for w status call wait. Does process
100 have any children? It has two children. Wait will return whenever the first one dies. What's the
one? Which one of these is going to be the first one that terminates? 102 because that's the one that got an
error, right? So 102 terminated first. Wait would return
pretty much instantly and say wait return for exit process in this case process 102 and it status would be
whatever the hell this value is which I guess is 10. So it would get the value 10 and then that's it. Then the parent
would terminate right? So that means my the one of the other children. So
101 would exist for two seconds. So when the parent exited, was it a zombie? Was it
an orphan? Was it both? What was it when the parent exited? >> Both.
Yeah. >> Just an orphan. >> Just an orphan. Right. So it's still
executing. It's still in sleep. It lasts for two seconds. The parent pretty much terminates immediately. So it is just an
orphan. It is not called exit yet. It lasts for two seconds. So it would got reparented up to let's just assume a net
and then whenever it finally terminated a net would have cleaned it up. So through this
I have created exactly one orphan. So if I go ahead compile that hopefully that is what exactly what happens. So what I
call this weight example. So yeah, I can see exactly what happened for from it. So here's my calling weight. They got
interchanged between them, but this is the weight that got an error. Oh, which I didn't do anything about. Uh oh, I
didn't even check check it. I didn't check that this is not negative Z one. So if weight pit is negative one, which
means we have an error. I'll just say return 35 or something. So I didn't even check the return value. So here yeah now
it works better. So calling weight calling weight from both that says wait return for process in this case it would
have been 102 and it would have been status 35 which would have been the error from this from process 102 is how
we name them. Then again 101 would have been an orphan got adopted by a net or something like that. All right, any
questions about that? That was fun. All right, well let's take our 10m minute break and then we will continue
on with even more ter more terrible terminology if you can believe it.
A zombie process occurs when a child process has terminated but its parent has not yet called the wait() system call to read its exit status. The process table entry remains allocated, wasting a process ID. Zombies cannot be killed or cleaned up until the parent acknowledges the termination.
An orphan process happens when a parent process terminates before its child. The kernel automatically reparents the orphan to the init process (PID 1) or the nearest subreaper. This ensures the orphan's eventual cleanup, as init will call wait() on it when it exits.
The wait() system call blocks the parent process until any of its child processes terminates. It returns the PID of the terminated child, along with its exit status encoded in an integer. This status should be decoded using macros like WIFEXITED() and WEXITSTATUS() to determine the child's exit code.
The init process (PID 1) is the ancestor of all processes in Linux. It has two critical roles: creating user-space processes at system startup and acting as a 'reaper' for orphaned processes. If init exits, the kernel panics because there is no fallback process to manage children.
The waitpid() system call allows a parent to wait for a specific child process by providing its PID. In contrast, wait() waits for any child to terminate. You can use waitpid() with a PID of -1 to mimic wait() behavior, waiting for any child.
A process enters the 'Z' state (zombie) after it terminates but before its parent reads its exit status. You can check a process's state by examining the /proc/<PID>/status file or using commands like ps aux which shows a 'Z' in the STAT column for zombie processes.
Execution errors or premature parent termination can create zombie and orphan processes. To manage them, always ensure parent processes call wait() or waitpid() to clean up children. For orphaned processes, init will reparent and clean them, but you can configure subreapers using prctl() to handle cleanup closer to the process tree.
Keep this summary
Save it to LunaNotes and it becomes a real note in your library — editable, searchable, and ready to turn into flashcards or a diagram. Free to start.
Save to LunaNotesOr summarise for another video.
This summary and transcript were automatically generated using AI with the Free YouTube Transcript Summary Tool by LunaNotes.
Related summaries
Comprehensive Guide to Operating Systems in 6 Hours for Semester Exams
This video provides a complete overview of Operating Systems, covering essential topics and exam-relevant questions. With 15 years of teaching experience, the presenter ensures that students grasp the core concepts effectively, making it ideal for beginners and those revising for competitive exams.
Understanding Operating System Design and Implementation
This lecture explores the complexities of operating system design and implementation, focusing on defining goals, user and system requirements, and the importance of separating mechanisms from policies. It also discusses the advantages of using higher-level programming languages for OS development.
Understanding System Calls: An Overview of User Mode and Kernel Mode
This lecture provides a comprehensive overview of system calls, explaining their role as an interface to operating system services. It discusses the differences between user mode and kernel mode, the process of context switching, and illustrates the concept of system calls through a practical example of copying file contents.
Linux OS Fundamentals and Shell Scripting Basics for DevOps Beginners
Explore the core concepts of Linux operating system and essential shell scripting commands in this beginner-friendly DevOps tutorial. Learn why Linux dominates production environments, understand its architecture, and master fundamental shell commands to navigate and manage your Linux systems efficiently.
Introduction to Operating Systems: Functions, Types, and Importance
This lecture provides a comprehensive introduction to operating systems, explaining their functions, types, and significance in computer science. It covers the role of operating systems as intermediaries between users and hardware, and highlights popular operating systems like Windows, Linux, and Android.
Most viewed summaries
A Comprehensive Guide to Using Stable Diffusion Forge UI
Explore the Stable Diffusion Forge UI, customizable settings, models, and more to enhance your image generation experience.
Kolonyalismo at Imperyalismo: Ang Kasaysayan ng Pagsakop sa Pilipinas
Tuklasin ang kasaysayan ng kolonyalismo at imperyalismo sa Pilipinas sa pamamagitan ni Ferdinand Magellan.
Mastering Inpainting with Stable Diffusion: Fix Mistakes and Enhance Your Images
Learn to fix mistakes and enhance images with Stable Diffusion's inpainting features effectively.
Pamamaraan at Patakarang Kolonyal ng mga Espanyol sa Pilipinas
Tuklasin ang mga pamamaraan at patakaran ng mga Espanyol sa Pilipinas, at ang epekto nito sa mga Pilipino.
How to Install and Configure Forge: A New Stable Diffusion Web UI
Learn to install and configure the new Forge web UI for Stable Diffusion, with tips on models and settings.
Found this summary useful?
Take it with you. One click puts it in your own LunaNotes library.
Save to LunaNotes