Course Overview
- Instructor: Bar Zini, Big Data lead at Mercedes-Benz with 10+ years experience
- Duration: 21 hours covering Tableau from beginner to advanced
- Unique approach: 250+ animated sketch notes simplifying complex Tableau concepts
- Includes free materials: datasets, Tableau sheets for concepts, calculations, visuals, and downloadable sketch notes
Key Learning Modules
Introduction to Tableau and Data Concepts
- Business intelligence, data visualization importance
- Big Data, IoT, data science fundamentals
- Tableau product suite overview: Desktop, Public, Prep, Server, Cloud, Reader, Mobile
- Tableau architecture: live vs extract connections, file types, server components
Environment Setup
- Download and install Tableau Public
- Create free Tableau Public account
- Use provided datasets (EU and non-EU versions) for practice
Data Modeling in Tableau
- Star schema fundamentals: fact and dimension tables
- Tableau data modeling layers: physical (joins, unions) and logical (relationships)
- Methods to combine tables: joins (inner, left, right, full), unions, relationships, data blending
- Practical creation of two data sources (small and big datasets)
Tableau Metadata
- Data types: number (integer, decimal), string, date, boolean
- Roles: dimensions vs measures
- Discrete vs continuous fields and their impact on filters and views
- Geographic and image roles
- Renaming conventions and techniques
- Aliases for data cleaning and abbreviation
Organizing Data
- Hierarchies: creating and navigating drill up/down
- Grouping dimension members: groups, clusters, sets, bins
- Practical grouping and clustering examples
Filtering Data
- Types of filters: extract, data source, context, dimension, measure, table calculation
- Filter order and impact on performance
- Sharing filters across worksheets
- Quick filter customization and best practices
- Sorting data: user controls and developer options
Tableau Parameters
- Creating dynamic, interactive dashboards
- Use cases: calculations, reference lines, filters, swapping dimensions/measures, dynamic titles, bins
- Parameter actions for user-driven interactivity
Tableau Actions
- Types: URL navigation, sheet navigation, filter, highlight, set value, parameter value
- Creating and configuring actions in worksheets and dashboards
- Best practices for triggers and user experience
Tableau Calculations
- Four types: row-level, aggregate, LOD expressions, table calculations
- 60+ functions including number, string, date, null, logical
- Nested calculations and syntax overview
- Practical examples for each calculation type
Chart Types (63+)
- Bar charts (row, column, side-by-side, stacked, 100% stacked, lollipop, bar-in-bar)
- Line charts (basic, multiple, dual axis, cumulative, difference, rank, sparkline, slobby)
- Pie and donut charts
- Tree maps and heat maps
- Bubble and stacked bubble charts
- Maps (filled, symbol, night vision)
- Histograms (single and dual measure)
- Calendar heat maps
- Waterfall, part-to-whole, correlation, ranking, distribution, spatial, flow charts
Dashboard Design
- Planning with sketches and container structures
- Vertical and horizontal containers, floating vs tiled
- Layout management and item hierarchy
- Adding content, spacing, formatting, coloring
- Filters and interactivity
- Navigation buttons and icons
- Final touches and testing
Real-World Tableau Project
- From user requirements to mockups
- Data source preparation and modeling
- Building charts and KPIs
- Dashboard assembly and formatting
- Adding filters, interactivity, and navigation
- Delivering professional dashboards
Conclusion
- Mastery of Tableau fundamentals and advanced features
- Ability to implement real-life BI projects
- Strong foundation for career growth in data visualization and analytics
This course is designed for beginners and experienced Tableau users, covering essential skills transferable to other BI tools like Power BI and Qlik. It emphasizes practical application, best practices, and performance optimization for effective data storytelling and decision-making.
For those interested in expanding their data visualization skills further, consider exploring Understanding the Weaknesses of Data Science and the Basics of Data Visualization for foundational insights. Additionally, if you're looking to enhance your skills in data preparation, check out the Comprehensive Guide to HR Data Preparation in Analytics. For a deeper dive into data analytics frameworks, Mastering HR Analytics: A Comprehensive Guide to Data Science Frameworks is an excellent resource.
hello and welcome to this very unique course to master Tableau my name is bar zini and I'm currently leading Big Data
projects at mercedesbenz with over a decade of experience in Big Data data visualizations and business intelligence
projects and I'm very excited to be your instructor for this course in this 21-hour course I'm going to be sharing
everything that I know about one of the most in demand skill in data science and data visualizations Tableau so that by
the end of the course you're going to be able to create amazing dashboard and visualizations in Tableau like I do in
the real projects and I designed this course to take you from Zero to Hero so if you are a beginner don't worry about
it I'm going to explain everything from the scratch step by step so that means this course assumes that you don't have
any skills in data visualizations and as well all the skills that you going to learn in this Tableau course like data
moding and so on could be used in any other tools like powerbi and click and now of course you might ask yourself
what makes this Tableau course different and unique from all other online courses this is the only course that breaks down
the complex concepts of Tableau into animated visuals because visuals are very powerful to make complex Concepts
easy to understand and to follow and in this Tableau course we're going to present over 250 animated sketch notes
of Tableau Concepts understanding the concepts and how Tableau work can make you a professional and expert in data
visualizations and in Tableau and in this course I'm going to provide you with tons of free materials like for
example I have prepared three different data sources for this course that we going to use in all our tasks and
examples through the course and as well I'm going to provide you with three Tableau sheet sheets one sheet sheet for
all Tableau Concepts another one for all Tableau calculations and we have one more sheet sheet for all the visuals to
help you choosing the right charts so having those three sheet sheets you don't have to memorize everything you
have a quick reference and access to Tableau Concepts and as well you have access to all Tableau files and
dashboards that is created during the course and as well all the sketch notes of each section are available to you to
download so you can use it later as a reference so now let's have a sneak peek about the Tableau course we will start
with the basics what is business intelligence data visualizations what is Tableau and then you're going to learn
the Tableau product Suites and after that we're going to do Deep dive into different Tableau Concepts like the
Tableau architecture Dimensions measures discret and continuous data after that we're going to Deep dive in tblo
calculations and fun functions you're going to learn more than 60 different functions in Tableau to manipulate your
data and after that we're going to go and cover more than 63 different types of charts in Tableau and then at the end
we're going to go and Implement Tableau projects similar to the one that I do in real life projects so now the question
is who is this course for if you are someone that has never built any data visualizations using tools like Tableau
or power pii I will be with you in this course in each step starting from the fundamentals and we're going to end up
having the advanced topics and this course is as well for you if you are already a tableau developer so I would
suggest for you that to take a look to the course curriculum and start at the level that suits you I have covered a
lot of advanced topics and you're going to have a lot of best practices in this course and this course is suitable for
you if you have experience in any other tools like in powerp and you would like to pick up a new skill in Tableau so
let's jump in and get started now we're going to have a quick overview of the tblo course I have
splitted this course into 15 different sections for example we're going to learn what is business intelligence what
is data visualizations what is Tableau and the history of Tableau and why Tableau is very powerful tool for data
visualizations after that we're going to go and deep dive into the Tableau product suit we don't have in Tableau
only one product we have eight different products so I'm going to go and introduce you to those products and
we're going to go and compare them side by side for you to understand the differences between them and I'm going
to help you to choose the right products for your project moving on we going go and deep dive into the Tableau
architecture here we're going to learn many different concepts like what is live and extract connections what are
the different types of Tableau files and then we're going to Deep dive into the Tableau architecture in order for you to
understand the main components of the architecture and how Tableau internally Works after all the theory we're going
to start preparing your environment in order for you to practice with me in this course so we will go and download
and install Tableau for free of course at your PC we're going to go and create free Public Accounts we're going to
download the training data sets and we going to publish our first visualization and at the end I'm going to take you in
a tour in order to make you familiar with the Tableau interface and after we have repaired your environment we're
going to start with the first topic how to create a data source in Tableau and here you're going to gain skills about
the data modeling so we're going to go through the basics of data modelling and as well how to do modeling in Tableau
and then we're going to go and learn four different methods on how to combine tables in Tableau using joints
relationships and data blending and of course we're going to go and compare them side by side for you in order to
understand the differences between them and when to use which method and at the end of this section we're going to go
and create two data sources moving on we're going to start talking about the Tableau metadata here you going to learn
very important Concepts in Tableau the data types dimensions and measures discrete and continuous values once you
understand those Concepts you can understand how to create visualizations in Tableau after this section we have a
small section about renaming so here we're going to talk about the naming conventions that each developer should
know then we're going to learn the different techniques on how to rename columns and tables in Tableau and at the
end we're going to learn how to give aliases to the values moving on to the next section you're going to learn how
to organize your data in Tableau and here we have different methods like grouping up the dimensions using
hierarchies grouping up the values using groups and clusters and then after that we're going to learn sets in Tableau and
at the end we're going to learn how to create pens in Tableau in order to create histograms and now in the next
section we're going to learn how to filter our data in Tableau and here you're going to learn the different
types and concepts of filters in Tableau how to create them and how to customize them and I'm going to give you 10 tips
and tricks about filters in Tableau and we will learn as well in this section how to sort our data after that we're
going to learn very important Concept in Tableau which is the Tableau parameters Tableau parameters are great in order to
add Dynamic to your visualizations so you're going to learn the concepts of parameters and then you're going to
learn different use cases for that how to make Dynamic calculations Dynamic reference line filters how to swap
measures and dimensions and to make as well Dynamic pens moving on to the next section we're going to learn as well
something about Dynamic so we're going to learn the Tableau actions in order to make your dashboards interactive so as
usual first you're going to understand the concepts of Tableau actions and then we're going to go through all Tableau
actions types for example how to go to URL how to go to Sheets how to filter data using actions and then how to make
highlights using actions and how to change the values of sets and parameters and after this section we're going to
have the Tableau calculations this section is very huge you're going to learn how to transform and manipulate
your data using four different Tableau calculations types so we have the RO level calculations aggregate calculation
table calculation and the LOD Expressions so in this section you you can learn more than 60 different Tableau
functions in order to manipulate your data moving on to the next section we have another big one we have the Tableau
charts here we're going to go and build together more than 63 different charts in Tableau so we will start with the
basic charts like the bar charts and we're going to end up building very Advanced charts in Tableau and at the
end I'm going to help you to choose the right charts for your requirements moving on to the next one we're going to
learn the Tableau dashboards we're going to go step by step on how to create clean dashboards in Tableau using
containers and now in the last section we have a tableau projects here in this section we're going to go together and
Implement a projects exactly like I do it in my real life projects so first we're going to learn the different faces
of each Tableau project then we're going to start with the requirements so you're going to learn how I analyze the
requirements of Tableau and then we start with the implementations of the projects so we're going to go and build
the data sources the charts and two different dashboards so with that you're going to get familiar on how to
implement projects and companies using Tableau so once you go through all those sections you're going to have a solid
knowledge about Tableau if you are new to the world of data you must start hearing a lot of
buzzwords from Big Data to iot data science data engineering and phrase it like data is the new oil in this
tutorial I will be covering some important buzzwords about the data and what they really mean so let's Dive In
we are living now in the data driven age and data is generated everywhere we people we generate massive amount of
data as we speak each click on the internet each search email or even if you are ordering something online we
generate data we spend hours every day on the social media liking commenting searching our smartphone is just all
time uploading data about where you are how fast you are moving and everything we do online is now stored and tracked
as data not only our smartphones and computers are connected to the internet and generates data but also we have
something called smart home we can connect any device at our home to the internet just put the word smart before
it we have Smart M smart lightning smart Fitness voice devices security systems all those devices could be connected to
the internet and start generating massive amount of data and this is what we call Internet of Things iot iot is
the concept of of connecting any device anything to the internet in order to generate and exchange data not only we
have iot at our home but also everywhere we are living in the digital transformation in the industry and
Manufacturing you might heard of the concept industry 4.0 the first Industrial Revolution introduced in
Germany it's all about smart factories connecting machines and devices to the internet in order to exchange data and
now we can find iots in the cities we are trying to implement those smart cities where we're going to connect
everything in order to reduce waste saving money improving quality we have as well iots in our cars our cars are
loaded with sensors and devices that are connected to exchange data for many reasons like driver assistance object
recognitions self-driving systems the list is just so long in 2022 we have around 14 billions of physical devices
things from small household cooking devices to the sophisticated industrial machines that are connected to the
internet generating and exchanging data the amount of generated data everyday is from iot social media
websites machines is truly mind-blowing there are currently over 44 zettabytes of data in the entire Digital Universe
that is 21 zeros so that means we are no longer dealing with normal traditional data we are dealing now with the big
data so what Big Data means there is three indicators that help us to understand whether our data is big and
are defined by the three V's the first V is volume well big data is Big with the growth of the internet mobile devices
social media iots the amount of generated data from those sources has grown dramatically the second V is
velocity in normal data processing we use to process slow data or we call it patch data once a day or something and
then we store it in the disk but in Big Data WS the sources are generating streams of data with very high speeds
that means we have to process and analyze the data in in realtime fashion and then we store it in memory instead
of disk and the third V is Variety in traditional systems most data types could be captured and Raw on structured
tables like database or excels but in the Big Data WS data often comes in semi-structured format for example
server logs in XML or websites or the data comes in unstructured format like videos audios images free text so in Big
Data we have not only to deal with structured data but also with semi-structured and unstructured data so
the Big Data terms means how we can efficiently store process and analyze our data when it has huge volume high
speed and different types in order to reveal significant values for the business but we still have a problem
with that all those generated data are raow data raow data are just unprocessed rows and rows of numbers that are really
hard to understand hard to read badly structured and almost has no value to the business almost 70% of the W's data
are unused raw data if left without processing and refining is just worthless waste of money waste of space
and it generate digital waste stores in very expensive data centers and that's why we have the very famous phrase of
the famous British mathematician Clive Hy data is the new oil well it means that we have to extract the row data
like we are extracting oil we have to refine it process it transform it into something useful and has value to the
business well what this really means is that most of the companies are sitting on very big field of a new oil raw data
and most of them understood that data is their most valuable asset they have to extract it they have to analyze it in
order to reveal Insight that could help them in order to make faster and better decisions and that's why most of the
companies are hiring Army of data workers as we know the demand for data scientist is increasing rapidly and the
supply is low so now what we can do with all those chaos all those generated unprocessed raw data well we can do the
following stuff so what we can do we can design or build a data architecture data
architecture is the process of creating blueprint on how we organize process and store our data into different layers for
different purposes so that architecture make it easier to manage protect and access our
data another thing that we can do with the row data is data engineering data engineering is very complex process of
Designing and building data pipelines and data storages in data engineerings we usually build ETL processes to
extract the row data from multiple sources then transform it and then load it to the Target storage in order to
make it highly available and usable for the data scientist or any other end user another thing that we can do is
data modeling so data modeling is the process of connecting the dots so what we're going to do is we're going to put
all the data into entities and objects then we describe the relationship between those entities in order to help
us and help the programs to understand how the data are related to each other another thing that we can do with
the row data is we can do data mining data mining is the process of analyzing massive amount of row data in order to
discover knowledge to discover business intelligence like patterns and Trends to solve problems and to mitigate
risks another use of the row data is that we can use it in machine learning in machine learning we are providing the
computers with two things first the RW and historical data together with the mathematical models and algorithms so
once the computer has those two things it's going to start training and practicing in order to perform tasks
like predictions so it's like human the more the machine practice and train the better and accurate the results going to
be and next we can do data science data science is the scientific study of data and it combines three major Powers the
power of programming languages together with the mathematics and statistics and the knowledge of specific domain in
order to uncover valuable knowledge and insights from our row data one more thing that we can use on
the row data and my favorite one is that we can use data visualizations so data visualizations is the process of
converting numbers and raw data which is normally hard to understand and to read into visuals and charts like Bars by
tree plots in order to make it easier to understand and easier to read which really helps in the decision making
there are many other things and processes that we can apply on the row data but these are the major fields of
work that we can use in order to convert the useless row data into knowledge that has significant impact of value to the
business [Music] all right let me tell you this story we
have shops in three different cities in Germany in stutgart we have shop Berlin and Hamburg and our three shops are
generating every business day a lot of raw data on Sales Inventory levels products staff cost and so on and now we
have group of people that are the decision makers like managers HR finance and they have many questions and
decisions to make so they might have questions for example what happened and another questions about what will happen
now if the managers try to find the answers from the row data they might find nothing and no answers because the
row data usually very complex and badly structures and they are really hard to understand and that's why they're going
to go and hire some data analyst for example in order to help them finding the answers from the row data so the
data analyst going to go and start analyzing the row data by doing some magic for example cleaning up the data
connecting objects together and aggregating the data in different levels and at the ends the result will be
communicated as for example spreadsheet to the decision makers and in the other hands the managers can hire data
scientist in order to help them finding answers about what going to happen or uncover unknown facts and insides so the
data science can as well go and start analyzing the raw data but this time using different methods like for example
data mining machine learning or train model in order to find new insights new knowledge and answers the question
question S at the end the output going to be communicated as well to the managers as numbers and spreadsheets now
both of the data scientist and the data analyst did amazing job working on the raw data and analyzing those stuff but
the problem here is that the output might be hard to understand and read Because those managers are usually
people that don't work directly with the data every day so this could lead to a big gap between those managers and the
results and now in order to bridge this Gap and make everything easier we can use the power of data visualizations and
the result presented from the data scientist and the data analyst should be converted from this boring numbers and
spreadsheets to visuals graphs and charts the visual representations of the data will just do the Magic by making
everything clear and easy and it's going to bring very easily the wow effect once you are presenting your results so it's
going to help the managers to immediately find their answers and they going to start making decisions using
the data this process we call it a business intelligence or as a shortcut bi so now the question is why
visualizations is so powerful with the simple Visual Communications you can make a huge difference since the start
of the humanity thousands years ago an early human use visuals in order to tell a story and until now in the modern age
the human still uses visuals in order to tell any story because we humans we are visual creatures we think in pictures
and in visuals if we see a tree our brain going to store it as a visual as an image in our brain studies say that
90% of the information transmitted to our brain is visual but if we read the word tree our brain has fair to
transform it to a visual before storing it which is way slower in fact the human brain processes visuals 60,000 times
faster than it takes more fact about our brain is that we remember most of what we see and interact with it's proven
then the human remember only 10% of things we hear and 20% about what we read and it's also proven that we
remember about 80% of what we see and interact with that's why we have the famous phrases of a picture is worth a
thousands words and seeing is believing having all those facts no wonder that in digital channels the visual content is
taking taking over posts tweets articles news presentations dashboards you can find visuals
everywhere so now the question is what is data visualizations or sometimes we call it data Vis data visualizations is
the process of converting boring numbers and raw data into interesting graphical elements like Bars by three blots and so
on so data visualizations brings the data to life makes you the master of Storytelling of the insights hidden
within your numbers so it's like an art of converting highly complex massive amount of data sets into something very
simple something very easy to understand and to interact with imagine yourself to be one of the managers and you have two
data analysts one of them is presenting the result in spreadsheet filled with numbers and the other data analyst is
presenting the result with visuals filled with graphical representations of the data and both are presenting the
same facts which report you will prefer I would go with the right one because the left one is just dry numbers pouring
and unlikely you will be able to spot any Trends and patterns so the main benefit of data
visualizations is telling a story which arms you with tools in order to make the right decision at the right time there
are many other benefits like seeing the big picture tracking Trends making smarter and faster decisions discovering
unknown facts acts patterns Trends and getting as well more engagement from the end users by asking more and better
questions all right so with that we have learned what is data visualizations and why it is very powerful and important
and next we will compare Excel to bi tools like Tableau and why you need to use Tableau instead of
excel over and over again I'm asked the same question why I should bother learning and using Tableau or power bi
for data visualizations if we have excel in this video I'm going to explain for you my six reasons why we should use a
more than bi tool like Tableau and Barbi and not use Excel for data visualizations and we start right now
there is around 1 billion users globally are using Microsoft Excel I worked in many companies and I can tell you people
are just addicted to excel they love it they use it for everything as planning tool data entry data analyzis and data
visualizations but the main problem here that the more a company grows the more it generates data and because everyone
is familiar with excels they going to keep using them in big data use cases and they're going to face really hard
time managing those spreadsheets and dealing with the limitations in Excel in these situations it's really time to
switch to a modern bi tool or data visualizations tool like Tableau or Barbi now let me show you how bi is done
with Excel we usually have different Source systems and data analyst that's going to go and start exporting manually
the data from those systems and import them in Excel and then some calculations going to be done and at the end a report
will be generated the Excel files then will be accessed from different business users in the other hand we can do bi
with a modern tool like Tableau so what we're going to do we're going to connect Tableau directly to those Source systems
and the data analyst can start developing a Rebo or dashboards in Tableau and at the end the business
users will access Tableau in order to see those dashboards so far you can say okay both look really similar so now
let's dive in in order to show you what is the real benefit of having a modern bi tool like tblo or Barbi and the
limitations that we have in spreadsheets like Excel the first benefit is automation if
you are using Excel and we made some nice reports it's time now to update the data and how we do that in Excel we
update data manually so some employee have to sit down every day and go through the process of extracting data
from those Source systems importing them in Excel do calculations and at the end prepare the reports over and over again
which is very time consuming but if you are working with the modern bi2 like Tableau we can automate this boring task
by creating schedule to refresh the data for example we can create a schedule in Tableau every day at 7:00 morning
Tableau should automatically connect to the data sources pulse the data and prepare the reports there is two
benefits of doing that first we eliminate the human errors which is very common thing in Excel and sometimes
those mistakes can lead to wrong decisions and to finance loss and the second benefit of course we no longer
needs employees that is dedicated only for this boring task of exporting and importing data manually to
excel another benefit here is the capacity if you are working with Excel and one of our source systems start
producing and generating massive amount of data here we have problem in Excel because we can handle around only 1
million records so our Excel file going to breaks and we're going to start getting error messages likes the data
set is too large so what we usually do in Excel we're going to go and start splitting the main file into small
multiple files in order to manage the huge volumes of data which is really hard to manage in the other hand if you
are working with Tableau we don't have to worry about all those stuff we have no problem in Tableau because Tableau is
made for big data use cases and can very easily handle massive amount of data so we might just change the connection type
from extract to live in order to handle it another benefit is security if you are working with Excel it's really hard
to hack into Excel even if you are using password protected spreadsheets it still can easily hacked nowadays and the users
are really used to share their excels in emails copy it USB or store it locally at their computers which is not secure
at all so all those stuffs could cost the companies a lot if sensitive and confidential data is accessed by
competitors but if you are working with modern bi2 like Tableau it going to provide us with Superior security
features like Advanced Access Control Data security network security and plus if you are working with Tableau we don't
have to export the data we can just share the dashboards and reports between employees and only if we grant them
access rights they can see the data another benefit is their role level security in many companies they have a
lot of confidential sources and they start to understand how important is to apply the principal need to know the
principles needs to know says a user shall only have access to the informations that's their job functions
requires that means we cannot go and share all data to all users we have to have some data restrictions for for
examples a sales employee should not see all data like manager and finance employee should not see all personal
informations like HR and so on that means if you are working with excels we have here again to split the main files
into specific reports for specific rule but in the other hands most of the modern bi tools they offer a feature
called Rowl security RLS row level security refers to restricting the roles of data as certain users can see based
on the policies that we Define using this technique going to enforce the need to know principle and going to make our
life easier by just having one dashboard accessed by different types of users and then based on their rule they going to
see the data and the informations that their job requires another benefit is reducing
chaos let me tell you how we usually work with Excel a data science will start exporting data from one source
system and he going to make a report called version one report and then for other requirements he going to make a
version two reports and eventually we're going to have a final reports and we have another data analyst working in
different Source system and the same thing going to keep happening few times back and forth and eventually we're
going to end up having different six versions of the reports and if we scale this impact you will notice that you are
slowly poisoning your business and the end user is going to have to access different versions of the reports and
now if we ask how old is the data in our reports we will get different answers when one version going to be 10 days ago
another one eight 4 and 3 days that's mean we don't have single point of Truth for our data and that's why having
modern bi tools going to help us to eliminate such a chaos and going to help us building a single point of Truth for
our data one last benefit that I would like to talk about is visuals although excels
offers visualizations but it is sometimes very limited when we are producing complex visuals in Excel is as
well creating visualiz Iz ation is very time consuming including a lot of manual steps and as well those visuals going to
be static and not interactive but in the other hand if we are using Tableau everything going to be automated and
super fast we can create new reports and Views very quickly by just drag and drop and they offer way more interactive and
cooler visuals than Excel all right the main reasons why I prefer working with Mod bi tools like
tblo and powerbi and not for data analyzes and data visualizations are automations security big data use cases
and interactive visuals it's not about Excel versus Tableau it's all about using the right tool for The Right Use
cases and not to misuse a tool Excel is a great tool that is used by billions of people because it's very easy to use
sheep professional spreadsheet for data entry and complex calculations but when it comes to data analyzis and data
visualizations we have way better tool than Excel like powerbi and Tableau and you can still use them together for
example you can do your complex calculations in Excel and the final result going to be imported in Tableau
in order to do better visualizations and to get more insight about the results the thing is the world is changing very
fast and the companies are generating massive amount of data so instead of using traditional spreadsheets like
Excel we have to use more powerful Tools in business intelligence to help us quickly find insights Trends patterns in
order to make faster and better decision so now the question is what are the best tools for data visualizations a
leading research company called Gartner publish every year the Gartner magic quadrants to show who are the leading
product in specific domain and if you check the magic quad for analytics and business intelligence platforms for the
last 10 years you can almost see always the same leaders we have Tau power pi and click view since 2003 12 and I'm
working with a lot of data visualizations tools and I can say that all those three tools are really great
tools they have their advantages and disadvantages but by just checking the data visualizations aspects I can say
that Tableau is here a winner because data visualizations in Tableau is a core concept and really the best tool for
data scientists and for Big Data all right the first question is what is tableau a quick answer could be Tableau
help us to convert this to this without any technical or programming skills so Tableau converts
complex and boring grow numbers into beautiful visuals and charts which is really easy to understand and the key
features in Tableau is interactivity easy to build and to use and fast performance we can call Tableau with
many names like a data visualization tool a business intelligence or bi tool or sometimes we call it a reporting tool
well Tableau is all of them but I choose to call Tableau a data visualization tool because data visualizations is the
core concept of Tableau now let's have a quick history about Tableau in 2003 tblo was founded
by three guys Pat Christian and Chris as a result of computer science project at Stanford University they focused in
visualizations technique to analyze data inside databases and then in 2019 Tableau was acquired by Salesforce in a
deal worth over 15 billion and for the last 10 years Tableau was named as a leader in Gartner magic cordance for
business intelligence Tableau has a clear mission to help people to see and understand
their data they really focus on keeping Tableau intuitive and easy to use that's why Tableau does not requires any
technical or programming skills in order to build amazing dashboards and insights that means the target audience of
Tableau is not only for technical users like it data analyst data scientists but also for all other non-technical users
like a business user an end user a teacher and so on this aspect is a GameChanger of changing the old mindset
of having only it and Technical people working with data and building visualizations but now we have modern
data visualizations tools like Tableau which opens the door for everybody to start start working with data that's why
tools like Tableau helps organizations to be data driven and now Tableau is widely used you can find Tableau almost
in all organizations Industries sectors in all departments because most of those organizations want to empower their
employees with tools like Tableau in order to make better faster and smarter decisions using data all
right tblo is not the only leader in business intelligence and data visualization Market there are many
other tools that are available like powerp click View and so on but now if you ask me what makes tblo so special
why Tableau is so widely used I would give you four reasons the first reason is performance
the sources now are generating massive amount of data and Tableau is designed and optimized to handle huge volumes of
data without impacting the performance in the dashboards and that's because is using high performance inmemory data
engine to help analyze large data sets where the data going to be stored inside columns instead of rows which can boost
the performance in dashboards tblo has no limitations or whatever to the number of data points in the visualization for
example on this view we have over 1 million data points without any problem this allows us to analyze large data
sets in order to find Trends patterns with a great performance and all other tools still inforce through size data
point limitations which is not really helpful for data analyzes the second reason is quick and
interactive visualizations compared to the other tools with Tableau we can create rich and beautiful visualizations
in just few seconds I'm going to show you now quick example how to Cluster my data and how to calculate the forecast
in order to do such a complex job in Tableau we will just use drag and drop so let's see how simple it is all right
so we're going to go to the order take the sales put it in the columns profit and the rows and take the order
IDs in the details and I want to see all my members over here and now we go to the analytics pan and then double click
on the Clusters so with that I have very nice four clusters of my data The Next Step I will create a forecast of my data
so I'm going to take the order ID put it in the columns and then we're going to take the sales I would like to change
the visual to bars so I have now here around five years what we're going to do we're going to go to analytics and just
click on the forecast and that's it so I have a forecast of two years of my sales and now I'm just going to go and put
them together in one dashboard so I'm going to create a new dashboard drag and drop the Clusters drag and drop the
forecasts and going to link them together with the filter and that's it so now we have both of them and if I
click around I will have an interactive dashboard for the forecast and for the Clusters
the third reason tblo is userfriendly as you can see we have done very complex analysis with just drag and drop without
writing any code and this is exactly what Tableau wants it's very intuitive and userfriendly and this is the major
strings of tblo it just opens the door for all nontechnical users to have a chance to work and play with data to
solve their daily problems without the need of it but in the other hand Tableau is integrated with programming languages
like python and R which opens another door for Advanced Data visualizations which might be used from data
scientists and the last reason is community if you are working with Tau well you are not alone you have a huge
Tableau community in the community we have around 2 million students and teachers and in Tableau public we have
around 5 million data visualizations that are published and there's around 200,000 questions and ideas that are
shared in Tableau forums having such a huge Community is a big Bloss for any tool it's very important because while
you are working with data you might face some problems or you have questions it's very important that you have a place
where you can go and ask your questions and get advices from other developers all over the world and not only that you
can as well get inspired from the shared visualizations from other developers you can find the important links about the
Tableau community in the video description below all right so my four reasons why Tableau
is one of the best tools for data visualizations are Tableau can handle massive amount of data very suitable for
big data use cases it offers beautiful quick interactive visualizations Tableau is intuitive and userfriendly no coding
or technical skills are required and the last reason table Community is very huge one more thing that I would like to add
that data visualizations is really one skill that you have to master as a data scientist or data analyst and Tableau is
an amazing tool for data visualizations that's why I highly recommend to learn or to get familiar with Tableau it's
going to be like a huge Advantage for your career all right guys so with that you know my reasons why I think Tableau
is a leader in data visualization and with that we have finished the first chapter of Tableau where we have covered
a lot of important terms of data and Tableau and in the next chapter we will have an overview of the Tableau product
Suite where I will introduce you to eight different Tableau products Tableau products in Tableau we have
eight different products and it's really important to understand them and understand the differences between them
so that's why I'm going to go and give you a quick overview of all eight table products and then we're going to go and
compare them side by side in order to understand the differences between them and at the end you're going to learn the
decision making process that I usually follow to choose the right product for your requirements so now let's start
with the first topic where we're going to have an overview of the the development process and products so now
let's go all right if you think Tableau is only one software then you are wrong if
you visit the homepage of Tableau tableau.com you will find many different Tableau products like Tableau desktop
public server cloud prep reader I can say other starts it might be confusing having all those Tableau products but
don't worry about it I'm going to explain them one by one so you can chose the right combinations of Tableau
products for you or for your organizations it's really important to understand the differences between them
the functionalities and the limitations of each Tableau products and let's dive in so Tableau product which contains
eight different products we have Tableau desktop Tableau public desktop prep server cloud public Cloud Reader and
Tableau mobile all right the first thing to understand is that we can split those products into two main categories
developer tools and sharing tools Tableau developers tools as the name implies they are tools that's going to
help you to build data visualizations by creating and designing dashboards charts reports or to do data preparations or
data engineering by preparing the data for data analyzis under this category we can find three Tableau products Tableau
desktop public desktop and Tableau prep and now in the other category we have the sharing tools those tools can help
you to share and collaborate your work work that you have done and created using the developer tools under this
category we can find five Tableau products Tableau server Tableau Cloud public Cloud Reader and tblo mobile all
right so now first let's focus on the Tableau products under the category developer tools and now we can go and as
well split the developers tools into two groups based on their purposes we have data visualizations and data engineering
underneath data visualizations we find two Tableau products Tableau desktop and Tableau public desktop and underneath
that engineering we have only one Tableau product and that's Tableau prep all right so now after we understood the
main categories and the main purposes of Tableau products we will go now and talk about the development process in
Tableau all right so basically we have three very simple steps in the development process in Tableau the first
step we connect our data to Tableau then in the next step we start building our data visualizations to do data anal izes
by creating report chart and dashboards and in the third step we share our work by publishing it the two products to do
these three steps are Tableau desktop and Tableau public desktop in many cases the quality of our data is bad and not
ready for analyzes that's why we add one more pre-processing step to prepare our data before we start building our
visuals and we can use for this step the product Tableau prep all right so now let's do deep Dives in into to tblo
developers products one by one in order to understand the key features and as well the limitations for each one of
them tblo disktop is a software you download and install at your PC with Tableau desktop you can connect to many
different Source types there are over 90 data connectors you can connect to Tableau server or to connect to files
like Excel text Json or to on Prem servers like my L and Oracle or to cloud like Amazon Google and Microsoft Azure
once you connect Tableau to your data you can start building your data visualizations in Tableau desktop you
will find many tools and functions to help you creating charts reports with just drag and drop and then you can
combine those different reports into interactive dashboards and after you're done building your views and dashboards
then you have three options to share your data by either publishing them to Tableau server Tableau cloud or to
Tableau public cloud or even you can store your workbooks locally at your PC all right so Tableau desktop is the
backbone product of Tableau as tblo developer you're going to spend 90% of your time using this tool tblo desktop
is a developer tool to build data visualizations where you connect your data build dashboards and then publish
them sadly tblo desktop is not a free tool like powerbi desktop in order to work with d desktop you have to buy a
license I think they offer some kind of trial phase or if you are a student you get get like one free year don't take my
words it's better to check the current offering from tblo in their homepage with tblo desktop you can connect over
90 different data sources you can publish as well your work everywhere to tblo server tblo cloud and tblo public
and since tblo discop requires a license you don't have any limitations or whatever on how many roads and data you
can store and process T desktop is meant for data analyst data scientist bi developers who work professionally in
companies in data anal iCal projects all right so that was a quick overview of the Tableau disktop next we will check
the Tableau public disktop so tblo public is the free version of tblo desktop it is very
similar to it it's a developer tool in order to build and publish data visualizations and since it's free and
requires no license it comes with few limitations in Tapo public we have a r 10 data connectors you can connect only
to local files at your PC another limitation of that you can store and process only 15 million rows of your
data and you can publish only to Tableau public Cloud so that means you cannot publish your work in Tableau server or
Tableau private cloud and the last limitation is that you cannot store your workbooks at your local PC but here I
have to be fair is that the most important part is that all functions and tools in order to build visuals and
dashboards are completely available in Tableau public like in Tableau desktop which makes really Tableau public as a
great alternative and tool for beginners in order to practice and to learn Tableau before they go and buy licenses
and to be honest that's why I decided to go with Tableau public in all my tutorials so that anyone can follow and
practice with me without having you buying any licenses tblo prep Builder is a software
you download and install at your PC and you can use it to prepare your data before you start analyzing it same as
tblo desktop you can connect to many different Source types there are over 90 data connectors like Tableau server
files on Prem cloud and so on once you connect Tableau to your data you can start building data flows where you have
access to tools and functions to help you to transform your data for example combining data cleaning filtering
aggregating and all other art of data engineering tasks to prepare your data for data visualizations and at the end
of your data flow you can store the new prepared data in three different places either as a file at your local PC or
publish it as a data source in Tableau server or cloud and the last option you can write the output directly in
databases and after you are done building the data flows then you can publish them in Tableau server or
Tableau online for automations and in Tableau prep you have the option to store your data flows locally at your PC
all right so TBB is a data engineering tool to prepare our data to get ready for analyzes sometimes the data that we
are connecting to Tableau desktop has bad quality and we cannot use it immediately in our dashboard that's why
we spend like hours and hours of cleaning up organizing combining preparing our data and that could be
really time consuming so for this situation we could use tblb to help us with this process so tblb is a developer
tool for data engineering where we connect to our data build data flows and then publish them and it's not free tool
it requires a license in t r we have over 90 different data connectors the output of the data flows could be stored
locally at your PC or as a Tableau Data Source or directly in the databases and we can publish the data flow either to
Tableau server or to Tableau Cloud tblo prep is not like tblo desktop we don't have any free version of Tableau prep so
there is no Tableau public prep all right so now let's go and have a summary of the three products where we
going to compare them side by side the main purpose of tblo desktop and public is to generate data visualizations but
the main task of tblo prep is for data engineering now if you are talking about the costs both desktop and prep requires
licenses but Tapo public is free to use and now about the security aspect of the data tblo desktop and prep are secure
since you can publish them to private servers but tblo public you have to publish your work to public platforms
where everyone can see your data so you cannot secure your data in Tableau public and the next Point data limits
since public is free it comes with the limitations of 15 Millions row but desktop and prep you will get no
limitations the next point is connectors in both desktop and prep you have over 90 different data connectors like files
API servers cloud and so on where in tblo public you can connect only to files and if we talk about the live
connections aspect the only tool offers live connections to your data sources is Tableau desktop you cannot make Live
Connections in Tableau public and in Tableau prep you have always to work with extracted data the next point is
about storing your files locally both tblo desktop and Bre allows you to do that by storing your work locally at
your PC but in Tableau public you cannot do that instead you have always to publish your work to tblo public Cloud
the last aspect is about the target audience tblo desktop is made for data scientist and data analysts but tblo
public is made for anybody wants to work with data visualizations and tblo prep is made for data
Engineers all right so now with this we have good overview of the three Tableau products for developments and now comes
at the question when to use which product so now let me guide you in my decision making process using the
following flu charts first we ask the question for which purpose if we need a product for data engineering then it's
easy we have only one Tableau product and that is Tableau prep now if we need a product for data visualizations then
we can ask more questions the next question do we need to connect to server API databases or to Cloud if the answer
is yes then we have to use Tableau desktop and if the answer is no then we ask the next question can our data be
public if the answer is no our data is confidential then we have to use Tableau desktop but if the answer is yes our
data can be public then we jump to the next question do our data sources contain more than 15 million rows if yes
then we have to choose Tableau disktop but if the answer is no our data sources have less than 50 million rows then we
jump to the last question do we need to have live connections to our data sources if the answer is yes then we
have again to choose tblo desktop but if the answer is no then finally we can go and use tblo public all right so if you
follow those questions and this chart you can easily decide when to use which Tableau
product all right guys so in the previous tutorial we splitted Tableau products into two main categories
developers tools and sharing tools now we're going to focus on the second category the sharing tools where we have
Tableau server cloud public Cloud Reader and tblo mobile and as the name implies those products can help us to share our
reports and dashboards with others and in the last last tutorial we have talked about the four steps of Tableau
development process now we're going to do Deep dive in the step number four where we're going to talk about the
different options that we have in order to share our reports and dashboards with others if you want to share your visuals
with your colleagues in your organization then we have here few options first you can install Tableau
server product on servers using the infrastructure of your organization and then you can start publishing and
sharing your dashboard there then your colleagues can either use their web browser or they can use tblo mobile app
on their smartphone or tablets to view and interact with your dashboards directly from the server the second
option we have we can install Tableau server products on cloud service providers like Amazon AWS Microsoft
Azure or Google cloud and then you can publish your dashboard there and the same thing here users can use web
browsers or Tableau Mobile in order to access your work the third option we have you can use Tableau private cloud
service here you don't have to install any Tableau server or anything you will get everything prepared from tblo Team
you can start immediately publishing your dashboard there and your users can consume it from tblo cloud and now let's
say you want to share your dashboards with everyone in the world and make it public then you can use tblo public
Cloud you don't have to install anything you can immediately publish your dashboard there and users all around the
world word can use their web browser to access your dashboards and data but they cannot use mobile app in order to access
Tableau public and now to the last option that I really don't like to use if you want to share your reports to
individual users you can send them a tableau file with the format twbx Tableau packaged workbook which
contains your data plus your reports and dashboards and then the users can view this file using Tableau reader software
installed at their p [Music] all right everyone so now in order to
understand the real differences between tblo server and tblo Cloud we have to understand the backend details and some
basic concepts about hosting servers let's go let's say we are startup company and we want to host our own
table application and build the entire infrastructure for that reason there is a long list of tasks that should be done
of course the first thing that you need to do is to go and buy some Hardwares and configure them like servers that
will run the applications and each servers need as well storage so we have to provide additionally storage
infrastructure like some hard disk driver and ssds servers needs to be as well connected to the internet therefore
we have to provide as well all the networking infrastructure once we have all those stuffs then we have all
Hardwares needed the next thing that we need to do is that we going to go and start installing and configuring some
softwares like we can install an operating system for example windows or Linux and many other middlewares once
the operating system is in place then we have to install and configure tblo server application once we have all
software and Hardware ready and running it's finally now the time to set up our Tableau project and we have to manage
the following tasks we have to start adding users to the Tableau server and map them to the correct licenses we have
as well to create schedules and tasks to refresh our data inside tableau server and then we have to start monitoring the
Tableau jobs all right so now we come to the big question that we have to answer who will manage what the first option
you have if you decide to manage all these layers that means we are talking about the on premises model so it's
clear ownership you manage everything from top to bottom Hardware the software and the project itself but now if you
say you know what this is too much to manage we don't have the money to buy all those stuff and hard Hardwares at
the start and we don't have the time to take care of them and maintain them then you will start thinking about
Outsourcing the Hardwares where you're going to buy a service from cloud providers like Microsoft Azure Amazon
AWS or Google Cloud so that they manage the hardware and you manage both software and projects and this is what
we call infrastructure as a service I the first letter of each word but now if you say you know what our it team is
very small we don't even have the time to keep those softwares updated each time Tableau makes a new release we have
to install a new version of Tableau server which is really wasting our time and we are not able to focus on our Core
Business project we don't have the resources to manage our own software then you start thinking about
Outsourcing the software layer to do that you can buy a service from Tableau it's called Tableau clouds where tblo
team going to manage everything for you both Hardwares and software and this is what we call software as a service
SAS okay guys so now let's summarize and compare the three hosting options the first point is about hosting setup on
premises you need Tableau server installed in your organization servers in as you need as well tblo server
installed in cloud service provider for example Microsoft Azure and in SAS you just buy tblo Cloud product and now for
the question who man what in on premises you manage everything the hardware software and your project and there is
no Outsourcing in is you manage both software and your project and the cloud service provider going to manage only
the hardware in SAS you manage only your business projects and Tablo going to manage both hardware and software so now
let's check the advantages and disadvantages of each service model for the on premises the good thing here is
that you have full control of everything the hard hardware and software and your data remains behind your firewalls this
is very important if you have critical or sensitive informations that should not stor outside of the company's
firewall but the drawbacks here you need a dedicated hardware and software administrators to deal with the
maintenance patching and many other tasks it is very costly at the start of the project you have to pay a lot for
the Hardwares and the softwares and it's not flexible it's really hard to scale up or skill down your Hardwares as
needed having all those stuff generally you have less time for your business projects all right so now let's move to
the is the First Advantage it gives you flexibility you can scale up scale down the Hardwares as the business needs and
there is no upfront cost for buying Hardwares but the downside of is is that you still need administrators to manage
your softwares to do installations patchings of your softwares and if you don't pay attentions for the cost you
might end up paying big bills now let's move to SAS the main advantage in SAS is that it allows your it team to focus
only on the Core Business projects and allows you to implement projects in very short time and the other good thing is
that your software will be always up to date Tableau team going to deal with that but the downside of SAS is loss of
control you will be at the mercy of Tableau team if anything bad happen like security problems all your
organization's data might be compromised and the other disadvantage is that you might have bad performance or networking
issues connecting Tableau to your Source systems and my advice here that you should avoid Reinventing the wheel
always take advantage of services that do things not part of your core business every hour you spend patching an OS or
installing updates for your software or replacing Hardwares is an hour not spent enhancing and refining your dashboards
in tableau [Music] all right everyone so now we're going to
do deep dives into Tableau sharing products one by one in order to understand their key features and as
well their limitations for each one of them and we start with tblo server and Tableau Cloud as Tableau developers in
organizations we need to share our reports and dashboards with the other colleagues in our organization so we
need to put those dashboards in a trusted environment or platform in our organizations and we usually have four
requirements the first requirement it should be safe and secure we want to control who is accessing our data and
dashboard second it should be easy to scale third it should be robust that can handles huge amount of users and data
and the last requirement it should be powerful and deliver high performance no one wants slow dashboards and reports
and now in order to build this trusted environment with these requirements we have two Tableau products Tableau server
and Tableau cloud and we have three hosting options on premises ASAS and SAS don't worry about the terms I'm going to
explain them Tableau server and Cloud they are very similar at the user interface level you will not notice any
differences but if you are checking the backend level there is a big differences between them so now first let's talk
about the user interface level of Tableau server and Tableau Cloud once you publish your dashboard to Tableau
server or Cloud you can share them by Prov find ing links to the users across all departments in your organization and
then the users they can access your dashboard using their web browser without installing any software at their
end and if you give them access they can start exploring your data in Tableau server or Cloud you can manage your
users by adding and removing them give them specific rules like admin creators viewers or Explorer you can manage your
users as well by adding them to groups another important task you can do in Tableau server or cloud is that so you
can automate your tasks for example you can create a refresh schedule to refresh your data sources on regular basis like
once a day in TBL server and Cloud you can monitor the tasks and schedules to check the status if the job failed or
succeeded and you can find many other statistics about the runtime the average queue and error messages and so on not
only the users can view the dashboards in Tableau server or Cloud but also they can create a new one if you give the
users enough rights they can even start creating their own insights and Views directly on their web browser without
having them to install any tblo desktop it's something we call self-service Pi all right everybody so now with this
we have clear picture about tblo server and Tableau Cloud so now let's talk about the other sharing Tableau products
Tableau public cloud is a free cloud service managed by Tableau team everyone in the world can share visualizations in
this platform so if you publish your dashboards in Tableau public everyone can access it interact with it and even
download it Tableau public is like social media you can edit your profile and add your personal informations in
Tableau public you have a huge gallery of visit built by people all around the world it hosts currently over 5 million
visualizations in Tableau public if you are browsing and you found some interesting dashboard like this amazing
dashboard from AAS you can add it to your favorites and then you can check what other visit did I just created and
published to public and like any other social media if you like her content you can go and follow her to see her new
updates and if you're inspired of one of her dashboards you can go and install the whole workbook to see how she did
build these amazing dashboards and see all details with that you are expanding the knowledge in Tableau development so
using Tableau public you can get inspired from others and you can get connected to other Tableau Developers
from all around the world and one more cool thing about Tableau public if you are searching for new job and you want
to flex your data visualization skills you can publish a lot of work in Tableau public and Link it in your CV so that
the companies can see how skilled are you in Tao so all these nice features makes Tableau public Cloud a very
attractive platform for sharing visualizations but now if you are talking about the security aspect it is
very limited the only thing that you can control is not allowed to download your visualizations or you can completely
hide it from others but you don't have any user access control like we have in tblo server or Cloud so Tableau public
cloud is a free cloud service from Tableau which host a lot of reports and dashboards built by people all around
the world it's a great platform to get inspired by Tableau Community build connections to other Tableau developers
and share your skills but since it's free it comes with few limitations the total size available for each account is
only GB your dashboard and reports are not connected to the source systems that means you cannot automatically refresh
your data in tblo public always you have to do it manually so you're going to open the reports refresh the data and
again publish it to tblo cloud and the third limitation of Tableau public is that as the name implies everyone in the
world can see and share your data that means you cannot use it in organizations since you cannot protect your data
[Music] Tableau reader is a software you download and install at your BC you can
use it only to view reports and dashboards but you cannot use tblo reader to create any data visualizations
or even edit it as you can see we don't have any tools or functions to create charts you can't even connect any data
sources or refresh your data Tableau reader is very old tool from Tableau it was created in the early days of Tableau
in order to share content build using Tableau desktop this was before even Tableau server and Tableau Cloud made
available at that time Tableau reader was the only option you have in order to share dashboard and Report with other
users so how it works you build data visualizations using Tableau desktop and then you send a file to someone else
then they going to use Tableau reader in order to view and interact with the dashboard that you built so to summarize
Tableau reader is a free tool it is just to view and interact with report and dashboard bu using Tableau desktop you
cannot create or edit anything in Tableau reader you cannot refresh the data inside your dashboard using table
reader each time you have to ask for a new copy if you want to have fresh data and there is no security features
password protections or login option this is big problem if the files lands on the wrong hand your organization data
could be exposed well I don't recommend at all using this tool in organizations the risk is just too big but if you want
to take take the risk and to share your visuals with one two three persons then use it but try to avoid
it tblo mobile is a free mobile app that you can download at your smartphone or your tablet you can use it to view and
interact with Tableau reports and dashboards published to Tableau server and clouds so you can use it only to
view the reports you cannot use it to create new reports or to edit the reports while tblo mobile is free to
download it requires a license to use and it can only access Tableau server and Tableau Cloud so you cannot use it
in order to access Tableau public and tblo Mobile going to automatically cash your reports and dashboards in memory
that means you can access them even if you are offline all right everybody so now let's
summarize and compare all Tableau sharing products side by side the first point about hosting Tableau server can
be hosted in your organizations or in cloud service providers like Azure or Amazon both tblo cloud and tblo public
Cloud are hosted by tblo team Tableau reader will just be software installed at your PC you can't even host it now if
you are talking about the cost for Tableau server you have to pay for licenses Hardwares and maintenance but
in Tableau Cloud you have only to pay for the licenses Tableau public and Tableau reader are free to use now if
you check the data security aspects both tblo server and tblo Cloud are highly secure Tableau public and reader they
are not next point is about the storage limitations in tblo server it really depends on the server dis space in tblo
cloud and reader there is no limitations but in tblo public Cloud the total size available for each account is only 10 GB
the next point about the connectors Tableau server and Cloud can be connected to different types of sources
like Cloud API Services files databases and so on but tblo public cloud and Tableau readers they cannot be connected
directly to any of your Source systems let's jump to the next Point Automation in Tableau server and Cloud you can
schedule tasks to refresh your data inside your dashboards automatically from The Source systems but the data
inside tblo public cloud and reader cannot be refreshed you have to do it manually you have to republish it or to
resend the file the next point about tblo mobile you can connect your smartphones or tablets only to Tableau
server or Tableau cloud and now to the last point we can use Tableau server and Cloud to share dashboards inside
organizations tblo public is used to share dashboards to the whole word and TBL reader is used to share dashboards
directly to individuals all right so now with this we have an overview of all Tableau
sharing products so now the question is when to use which product so let me guide you in my decision making process
following this chart all right first we ask all questions about the limitations inside tblo public Cloud the first
question can data be public if the answer is yes then we ask the next question should the data be frequently
refreshed in the reports and dashboards if the answer is no then you can go and use tblo public clouds but if the data
should not be public and should be refreshed automatically then we have to think about private hosting for the
question now do you want to manage the hardware if yes then you can use Tableau server on on premises as your
organization but if you don't want to do that and you want to Outsource it then you ask the next question do you want to
manage the software on your own but if the answer is yes then you can use again tblo server but this time it's going to
be hosted in cloud service provider like Microsoft Azure in is service model but if the answer is no you don't want to
manage the software by yourself and you want to Outsource it then you can go and use tblo cloud as a SAS Service as you
can see TBL reader is not in my decision making process since I don't recommend it at all so now if you combine this
flowchart with the one that we built previously for developers tools you will get my whole decision making process
that I usually use when I start a new Tableau projects so if somebody asked you when to use which tblo product you
can go through it and find the right combinations for you or for your company all those materials you can find it in
my website all right everyone so with that we have covered all eight Tableau products and we understood the
differences between between them in the next chapter we will learn the Tableau architecture to understand how Tableau
internally works and what are the main components of Tableau Tableau architecture now we're
going to go and understand how Tableau internally Works its components and its limitations so now we're going to go and
cover many important Tableau Concepts like what is live and extract connections what are the different file
types in Tableau and then we going to start drawing the Tableau desktop architecture and then we're going to
jump to Tableau server in order to understand different scenarios like the publish process authentication process
and accessing view process and after that we're going to go and complete the big picture by drawing the server
architecture and its components and at the end we're going to cover as well the architecture of the Tableau public so
now let's start with the first concept the live and extract data connections so now let's
go how we come to the most important decision or questions that's we're going to make make inside data source do you
want to store an extra copy of your data inside Tableau so here we have two designs for the data source either
you're going to say no we don't need to copy inside Tableau the data should stay where it is in the source systems then
what going to happens each time your visualizations needs data it going to sends quaries directly to the external
database and then the database going to send the results back to your visualizations so the data comes always
fresh from the s sources directly to your dashboards this type of the connections we call it a live connection
or you're going to say yes let's have a copy of our data inside Tableau so a snapshot or subset of the data going to
be copied from the external database to Tableau this copy we call it an extract and now each time our visualizations
needs data it going to send queries this time to the extract instead of the external database and then the extract
going to return the results back to your visualizations and since the extract is inside Tableau and very close to the
visualizations we will get great respond time and very fast performance this type of connection we call it an extract
connection all right so now the question is which connection type should I use in my data sources the typical answer for
this question is well it depends because here we have a tradeoff between performance and data freshness for
example if for you the performance is way more important than the data freshness then you have to go with the
extract since the dataa going to be stored inside Tableau in memory using the column store technique you will get
just great performance but if you say you know what the data freshness for me is more important than the performance
then you have to go with the Live Connections in your data sources because you will always get fresh data directly
from the sources in your dashboards all right so now if you want to send Tableau files directly to the
users we have to ask ask the question which type of files we're going to send because in tblo desktop we can generate
not only one file we can generate five different types of files in Tableau so now we're going to have like quick
overview of those types of files to understand them and to know when to use them all right so as we learned the
Tableau workbook contains three things the extract the data source and the visualizations there is a file type for
each combinations depend on your requirements for example if you want to share only your data without anything
else no data source no visualizations then you can send an extract as a hyper format but now if you say you know what
I've done a lot of work in the data source I built a data model I renamed stuff I did aggregations I created a lot
of new columns so I would like to share that with my team with my colleagues and I'm not allowed to share my data with
them so in this situation you say okay I'm going to share the data source with my colleagues and we call it Tableau
Data Source TDS without data or you might be in other situation where you say you know what my colleagues don't
have an access to the source systems we cannot use the live connection and you don't mind sharing your data as well so
now you can send them a package of an extract and a data source so the file type here called tblo package data
source DDS X so this type of file contains both of your data and your data source and we might be in another
situation where our colleagues or users they are interested as well in the visualizations so we can send them a
file with the visualizations and the data source and here again we have the same situation you decide whether you
going to send with it a data or not so if you don't want to send the data inside it you can send a file called
Tableau workbook TW WB and the last scenario I think you already guessed it if you want to send everything the whole
package the extract the data source and your visualizations then you can go and send your colleagues a tableau format
code tblo packaged workbook twbx all right so as you can see Tableau
did design different types of files for different purposes so depend on the situation or the scenario that you have
you can share your work with your colleagues all right so now generally speaking we have two different types of
workbooks a workbook with data using extract connection and another book without data using live Connection in
one hand in the workbook with data you can send three different types of files you can send only the data using hyper
format or send the whole data set with the data using TDS X format or send the whole package with the format TW wbx and
in the other hand with the workbook without data you can send only two files the data set without data TDS or the
workbook wbx and now you might have the question and you say okay which Tableau products should I use in order to open
these Tableau files well we have three Tableau products tableau Des toop Tableau public and Tableau reader with
the Tableau desktop you can open everything you can open all these different Tableau formats and files but
with the Tableau reader and public you can open only the Tableau packaged workbook TW wpx since Tableau reader and
Tableau public cannot connect directly to the data sources and they cannot use the live
connections all right one more thing to understand about Tableau workbooks is that Tableau uses two different types of
data to store the workbook the first one is the metadata information it will be stored in XML files metadata is data
about your data it describes your data it contains all informations on what have you done in the workbooks anything
you click drag and drop or do while working with dblo desktop will be reflected in some way in the metadata
you can find informations for example like column names data type data model and so one and the second type is the
data itself the actual data if you load data inside Tableau Tableau going to store it in format of hyber file where
the data going to be stored in column store methods in the memory of Tableau it is like special formats for fast data
retrieval all right if you understand the Tableau architectures and how the components are connected to each others
everything going to make sense for you as you are working with t and as well it going to makes you a
better Tableau developer so I will be sketching the concepts in order to make it easier for you to understand so let's
go the TBL architectures contains four different layers The Source layer the disktop layer server layer and the
consumer layer we will start unboxing each layer one by one to understand their components and we're going to work
with this architecture from left to right so we will start by The Source layer and we're going to end up by the
consumer layer all right so now we have the source layer The Source layer is outside of
Tableau and it contains the source of our data so our data could be in databases like MySQL or Oracle or the
data could be in files like Excel and Json or even in the cloud like Amazon AWS or Microsoft Azure or even in epis
so our data could be everywhere all right so now back to the big picture let's jump to the next layer
we're going to unbox the desktop layer the first component in tblo desktop is the data source before you start
building your visualizations you must set up the data source the first thing that we're going to do inside the data
source is to connect Tableau to our data Tableau offers around 90 different data connectors so we can connect Tableau
almost to anything once you build the connection between Tableau and your source of data the access information is
going to be stored inside the data source for example the path of the file location of servers username passwords
or access tokens and so on so all these informations going to be stored inside the data source all right so the two
types of data Connections in data sources are extract and live connections so now we connected to data we decided
which type of the connection the next thing that we have to do in the data source is to start building our data
model and we can do that by combining tables together using relationships joins and Union and you can do many
other stuffs like setting the right data types doing aggregations renaming stables and columns creating new
calculations and filters and so on all right so now to summarize the data source component in Tableau contains the
following informations we have the data connectors to connect Tableau to our data we have the access informations
where the locations of our sources going to be stored and as well we can decide whether we're going to load an extra
copy of our data inside Tableau we call it an extract connection or we're going to leave it as Live Connections in the
data sources and the last thing we have the data model inside data sources where we can combine tables together and do
aggregations or we can do some other custom stuff all right so once we are done with the setup of the data source
we have the connection whether it's extract or live we have our data model and everything is ready now we're going
to go and start building our visualizations and Tableau organizes the visualizations in three levels the first
one is the worksheets so we can use the data available in our data sources to build a single view only one visual it
could be a bar charts a pie charts or a table View and as you can see each worksheet is connected directly to a
data source but in Tableau you can build a worksheet from two different data sources by using very powerful combining
methods called Data blending and this is very unique feature in Tableau you cannot find it in any other bi tools
where the data in one visual can come from different sources and once we have the these different worksheets we can go
to the next level where we start combining these worksheets into One dashboard to show the different visuals
in only one view but keep in mind if you want to do any changes in the visuals you have to go back to the worksheets
and do the adjustment there and now we come to the last level we have the stories as you know the main goal of
doing data visualizations is to tell a story so you can build like a sequence of worksheets or dashboards that works
together in order to tell the users story based on your data all right so now you might ask me which visualization
level is the right one for you well if you have only one visual then go with the worksheet but if you want to build
some kind of qbi to monitor process then build a dashboard and if you want to present your data and tell a story from
it then go and build a story all right so now we have in Tableau desktop both of the data sources and the
visualizations and these two components are contained in something called a tableau workbook so now the question is
after you're done building your data sources and visualizations what can you do with this workbook well you can share
it with your colleagues in your team or department and there is two ways to do that either you're going to go and send
a tableau file directly to the users or you're going to go and publish the workbook to a tableau server or cloud
and from there your users and your team can access your workbook all right so now back to the big picture
the Tableau architecture let's talk about the ler on the right side the consumer layer there is different ways
to consume Tableau visualizations depends on the users's clients and on the tasks the users do so we start with
very small group of users that they might use Tableau reader to view and interact with Tableau visualizations and
they usually don't want to edit or create something new for this group of users we're going to send them a TBL
file as we learned they're going to need the Tableau packaged workbook twbx and we might have another group of
users usually they are your team colleagues they want to build analyzes on top of your work they're going to use
dblo disktop to do that for them we can send any kind of Tableau files depends on their requirements and their tasks
and now we have a big group of users or consumers that they can access Tableau server or Cloud to view and interact
with Tableau visuals they can use their web browsers like Google Chrome and Firefox to access the content of tblo
server and from there they can view interact and even edit the visualizations if they have enough
permissions or they can use Tableau mobile app on their smartphones or tablets to view and interact with your
workbooks but they cannot use it in order to edit the table visualizations so for this group of users you will not
send them any files first you have to publish your work to the server and here we have two options either you're going
to publish only the data source or you're going to publish the whole workbook to the Tableau server or cloud
and after that you're going to share the link of your workbooks to the users and now to the last group of users that's
worth mentioning they are the static users you can always export your data and visuals from Tableau desktop and
send it directly to the users as a BDF or Excel so of course it's static and they cannot interact with it all right
so so far in the Tableau architecture we talked about the source layer we did Deep dive in the Tableau desktop and its
components and we understood the different type of consumers and the clients all right so previously we start
sketching the Tableau architecture where we learned about the source layer the desktop layer and the consumer layer now
we going to unbox the server layer in Tableau architecture and in order to better understand Tableau server
components I'm going to walk you through three three scenarios from the user point of view what can to happen exactly
in tblo server once we publish a workbook or when we log into the server and access a workbook so let's
go so let's say that you want to publish a tableau workbook with an extract what going to happen tblo desktop going to
request the server to upload the workbook twbx and the first component in TBL server that going to receive the
request is the Gateway the Gateway knows how to forward the request the right server components and in this situation
the right component to process the publishing is the application server so the Gateway going to forward the request
to it and as we learned the Tableau workbook holds two different types of informations the metadata stored in the
XML files and the data itself stored in hyper files and in Tableau server those two different types of files going to be
stored in two different places the application server going to send the XML file to be stored in the server
component called repository and the Hy file going to be stored in another component called the file store so what
we have learned so far the Gateway is responsible to forward the request to the right component the application
server is the one that going to handle the publish process the reposter going to store the XML files the metadata of
the workbook and the actual data the hyber going to be stored inside the file store all right so now our workbook and
our data are published to Tableau server it's it's time now for our users to log the Tableau server and start interacting
with our dashboards so let's see how this going to work let's say your manager is Michael Scott and Michael
wants to check your sales dashboards in Tableau server and I am going to do it I need a
username and I have a great one and once Michael give these informations a request going to be sent
to the server as HTTP request the first thing that's going to hit is the Gateway the gateways notes that the application
saver is the right component to handle the authentication process so the Gateway going to forward it to it and
then the application server going to ask the reposter to check if the credentials username and password are correct and if
Michael has permission to access our server and then the reposter going to check and if everything matches and
Michael is allowed to access our server it will respond back to the application server and going to say yeah we know the
guy he is in our records then the application server going to start building the server UI and send it back
to the Gateway and then the Gateway going to send it back to Michael browser and now he is inside our Tableau server
so what we have just learned from this process again the Gateway is responsible of forwarding the request to the right
component the application server is the one that going to handles the authentication process the reposter
going to store the user credentials and if the users have an access and permissions to our server and the
application server is the one that renders the web interface of the server all right so now Michael is
inside our tblo server and he going to start browsing and searching for your sales dashboard and once he find it he
going to click on it and try to access your dashboard so now let's see what going to happen in Tableau server and as
usual the HTTP request for accessing going to be generated and sent to the server and we know by now that the
Gateway going to receive the request and start forwarding it to the right component the application server and
then application server going to start render the Chrome around the viz all those icons and images that are not
inside the dashboard itself and then the application server going to say okay now we are talking about visualizations this
is completely out of my league we have to forward this request to the master to the brain it is the vql server it is the
one that deals with visualizations and from here the vql going to take over and going to say okay first thing first
let's check if this guy Michael is allowed to see the sales dashboard so the vql going to ask the repository and
in the repository there is a list of users and reports so it's going to search there to find any matches if yes
then it's going to send back yeah Michael is AOS and he's allowed to see the sales dashboard and now the vql
going to say all right now we need data so first we need the metadata of the dashboard and as you know after we
publish the workbook the metadata going to be stored inside the reposer so the visq going to request from the reposer
one more thing is to send the XML file of the dashboard the reposter then going to send back the XML to the vql server
and the server will start building the dashboard all right so now the vsql going to say okay now we have the
dashboard but the problem is it is empty we need the data to fill it and it's better to ask our data specialist and
that is the data server the data server is the one that knows everything about the data so it's going to say all right
for this dashboard part of the data we have it already inside tblo server but the other part is sadly outside of
Tableau to get the data inside Tableau server from the extract the data server going to send the query request to the
data engine and the data engines knows how to query and extract the needed data from the file store so the data engines
going to get the data from the file store and it's going to send it back to the data server and now we come to the
part where the data is living outside of Tableau server here the data server going to act as a proxy where it's going
to use the data connectors to connect to the external data databases once the connection is established it going to
send a query that matches the language that the database speaks and then the database going to R return the needed
data as a raw table and now once we have all the needed data inside the data server it's going to combine it and do
another Security check so the data server going to check is Michael allowed to see all data or should we filter the
data so the data server going to filters the data depends on the data security setup that you have made and then it
going to send the raw data back to the the vql server and now once vql server has the raw data for the dashboard it's
going to do now the Magic by turning all those numbers and row data into images and visuals and it's going to put it
inside the workbook so now finally the vql has everything it needs the sales dashboard is complete and ready so the
vql going to send it back to the Gateway and the Gateway going to send it back to the web browser of Michael so Michael
can start interacting with the dashboard and now will hm does Michael have any idea what to do with the sales
dashboard I declare Banky all right I know there was a lot
of stuff going around in this scenario but we have covered most of the Tableau server components so let's have a
summary and understand what we have learned so far as usual the Gateway is responsible to forward the request to
the right component the application server is not responsible for the visualization process but the visl
server is the one that is responsible of building the visualizations the repository going to store informations
about the permissions and security which users are allowed to access which dashboard and the data server going to
manage both of the extract and live data sources and the data engine is responsible of retrieving the data from
the extract inside Tableau and the data connector is going to help the data server to connect to the external
sources and the vql server does the magic of transforming the raw data into visuals all right so so far with those
three scenarios we covered the most important comp component of Tableau server in this video you will learn
about the Tableau server architecture and then we're going to do a deep dive into each server component of the
architecture to understand how it works and what it does and we start right now the server layer contain mainly of three
stuff two interfaces left and right and in the middle we have bunch of server components the left interface is the
data connectors they going to connect the external Source systems to tblo server components and in the right side
we have the Gateway it going to receive requests from different clients and it going to connect it to Tableau server
components all right so now let's go more in details about the gate component so in one hand we have requests come
from different clients like a login request from web browser or a publish request from Tableau desktop and in the
other hand we have different Tableau server components like the app server vql server and so on and the Gateway
going to be in the middle that knows how to forward the requests from different clients to the right server components
and the other task of the Gateway is balancing stuff around let's say that you are working in multi- Noe
environment where you have two nodes when the Gateway receives the first request it going to forward it to the
node number one since both nodes are free but now if the Gateway gets a second request it going to say oh node
one is full let's process this request in node number two since it's free and so on all right so the Gateway in
Tableau server is like a distributor that knows everything you know something someone like that let's just say I know
a guy who knows a guy who knows another guy so the Gateway has two tasks first it Roots the client requests to the
right component and second it does load balancing if you are running tblo server in distributed environment all right so
now we're going to start talking about those Tableau components in the middle and in Tableau server there's like
different Arts of components we have servers we have engines and storages and we're going to start with the servers as
you learned in tblo server there is like different processes that login process publish accessing workbook and so on and
in tblo server they designed different servers for different processes so let's start now with the application server
the application server is responsible for different processes like as we learned a user login request going to be
forwarded to the application server then the application server going to check with a repository or an active directory
depend on your configurations to find out if the user is allowed to access the server or not and the other process the
application server handles is the publish process where the application server going to get the publish request
and it's going to split the workbook into two files the XML file to be stored in the repository and the hyper file to
be stored in the file store and one more task for the application server is to render the server interface all those
little stuff that you find in Tableau server like icons images projects minus it is the application server who render
those stuff so the application server is responsible for different processes like the authentication and authorization
process the publish process and rendering the server UI but one process that the application server will never
do is the visualization process all right so now we're going to jump to the next server we have the vql server this
one going to be interesting all right so previously we talked about the power of visuals and how human brain transform
text into visuals and images the vql is like our brain it can ad do the Magic by converting numbers and text into Visual
and images this quill stands for visual query language for databases the found of Tableau Chris and bat they did invent
this language let's say that you drag and drop something in Tableau the vsql going to convert this action to an SQL
query and then send it to the data server to get the data then the data server going to send the results back to
the vql as raw data and now vql going to do the Magic by converting those raw data into visuals and images presented
at your client all right so the vql is the brain it is very important Tableau component and responsible of the
visualization process and mainly it does two things it's going to generate queries from user action and it's going
to convert and transform the row data into visuals and images all right everyone so now we're going to talk
about the third one we have the data server so the data server is the one that knows everything about the data it
knows where to find the data how to connect to it and how to speak to it the first task of the data server is to
manage both extract and life data sources if the data is inside Tableau it going to send query request to the data
engine but if the data is outside Tableau it can to use the data connectors to send query request to the
external sources and the data server knows how to speak to the sources it act like a proxy to the data sources can
speak many different database languages so that it send a query request in a language that the database understands
and we have another task for the data server is to handle the data security it checks if a user is allowed to see the
data and do filterings if needed and the data server manages as well the driver deployment so the data server is the
central data management component in Tableau server and the one that knows how to get data from the sources all
right so now let's jump to the next component we have the data engine if we decide to store our data inside Tableau
as an extract then the data engine going to be the one dealing with it different components can send request to the data
engine like for example the data engine can receive a request from application server to publish a new extract then the
data engine can execute and create operation to create a new extract and store data inside it the data engine can
receive as well a query request from the data server asking for data so what going to happen here the data engine
going to find the correct extract it's going to connect to the hard driver and then it going to pulse the needed
extract from it and at the end the data going to be sent back to the server and finally the data engine can receive a
request from the backgrounder to update the content of an extract so the data engine can execute an update operation
by opening the extract and updating its content with the new data so the data engine in Tableau is like any other
database engine it does different operations like it queries the data it perform insert and update operations and
it create new extracts but only for the data inside Tableau server inside the extracts okay the next component is the
reposter as you might already noticed the reposter was involved in every tablet process so let's talk about it
the reposter stores many different types of data like for example it can store the workbooks that we published to the
server but only the metadata parts not the data itself so the XML files from the workbooks can be stored inside the
reposter in the reposter we find as well the usage data it's data that going to help you to understand the performance
and the traffic about your project like for example you can find the total number of active users inside tblo
server what is the total view counts by day and you can find out the most used data sources in your projects another
type of data that you're going to find inside the reposter is the security informations for example which users are
allowed to access your content or which users are allowed to access our Tableau server all right so as you can see in
the reposter there is different types of data and it contains as well huge amount of data in tblo server but it's very
important to understand that is the data inside our dashboards and reports are not stored inside the repository we have
many other Tableau server components that's worth mentioning like for example the cash server it stores almost
everything like images icons results of queries dashboards and so on so if you start a dashboard that that is already
accessed before the data going to be pulled from the cache server another component is the backgrounder in Tableau
server you can create a schedule to refresh the data inside your extract and the task of the backgrounder is to check
this schedule each 10 seconds and then trigger the process of refreshing the extract if the time comes and the last
component that I would like to mention here is the search and browse the users of Tableau server they can search for
content and this component is responsible for searching inside the repository and return the result to the
users all right everyone so finally we have the last puzzle the server components if we put it in the
architecture we will get the whole big picture of Tableau architecture so now let's go and do very quick summary that
Source layer it is the one that is outside Tableau and contains our data and it could be anywhere like databases
or files in the desktop layer the developers can start connecting Tableau desktop to the data sources with either
copying the data inside Tableau using an extract connection or with a live connections to the sources then the
developers can start building visualizations using worksheets dashboards and stories and both of the
data source and the visualizations we call it a workbook and we can either send it as a file or share it to the
server the server layer going to host our workbooks and we can find many components like the data connectors to
connect our sources to the Tableau server and the gateway to connect the client request to the Tableau server and
we have the application server responsible for the login and Publishing process processes the vsql server
responsible for the visualization process and the data server is the one responsible for the data management and
we have another component like the data engine that's going to handle the extracts and in Tableau server we have
three places where the data going to be stored and we have the reposter that contains many different data like the
XML of the workbooks and the security objects but not the data itself because our data going to be stored inside the
file store as an extract and we have the cache server that contains many different types of data to increase the
Tableau performance and the last last one is the consumer layer here we found the different groups of users and
clients like the Tableau readers that needs only the twbx files directly from the Tableau developers and another group
of users that they're going to use Tableau desktop to develop new views and we have the static readers that's going
to receive files like PDF and Excel and then we have a big group of users that's going to access Tableau server using
either web or Tableau mobile to interact with the published workbook all right everyone so one more
thing that I would like to show you is this amazing dashboard from Tableau team it's going to show you the different
component inside Tableau server and how they going to interact to do a task so for example if we go to the workflow or
to the process we can select for example access View and then we can to select whether it's like an published extract
or live and over here we have like slider if you drag it to the end you're going to see how the components are
interacting with each others to do the tasks and and on the right side you will see description for each step and this
is really great way to learn how tblo server works I learned from this a lot for this tutorial so make sure to check
that if you want to see more details about other processes in Tableau server I'm going to leave the link in the
tutorial materials let's start with the source of our data in tblo public you can only
connect files like CSV Json Microsoft Access and Google Sheets the next component is is Tableau public desktop
it is free version of Tableau desktop it's software that you can download and install at your PC so here we start by
connecting Tableau public to our files by creating a data source and in the data source we have only one type of
connection it is the extract so the data should be copied from our files to be loaded inside Tapo public desktop so
there is no live connection option and then after that we're going to start building our visualizations or we call
it visit and now once we are done building the views and the dashboards using tblo public desktop we have here
only one option to share it and that is to share the whole workbook your data and the visas to Tableau public and tblo
public is a free platform hosted from Tableau team to share the visualizations from the whole world and once our visits
are published to Tau public they can be now consumed from users all around the world and here we have few options the
users can use their web browsers to view and interact with your visualizations or users can download the whole workbook
your data and the vises in different formats like Tableau file twbx or Excel PDF images and so on and the last option
of consuming your visas can be embedded into your websites and blogs okay so now since Tableau public
is free it comes with few limitations at the source level we can connect Tableau public only to files the data connectors
are very limited and we cannot connect for example to servers and in the next level at the public desktop level there
is limitation in the data source we have only one type of connections and that is the extract so we cannot have a live
connections to the sources and the workbook itself it can contains only maximum 15 million rows and we cannot
save the workbook locally at our computer the only option to share it is to publish it to the Tableau public but
there is like a work around for that I'm going to show that in the next tutorial all right so now let's move to the
sharing level to Tableau public here we have as well few limitations for example the total available size for each
account is only 10 GB and there is no way to refresh your data automatically each time you need new data you have to
manually republish the workbook with new data and the third one it's going to be public so there is no way to make it
like private and to share it with only few people you have always to publish it to the whole world and now let's move to
the final level we have the consumers the only limitation here is that you cannot use Tableau mobile to access and
interact with the visualizations all right everyone so I decided to use tblo public in this Tableau course since it's
free and all of you can follow me with the examples without having you to pay for extra licenses and the limitations
that we have in Tableau public they are not really relevant for the learning process so the main features of Tableau
the data visualizations that we have in Tableau desktop they are all available as well in tblo public without any
limitations so don't worry about it all right everyone so with that we have learned the Tableau architecture and its
components and we learned how Tableau internally works and with that we have covered the theory parts of Tableau and
in the next section we will start preparing your environment so you can practice Tableau with me during the
course so let's jump [Music] in all right so let's start with the
first step we're going to go and download Tableau public desktop so in order to do that we're going to go to
the website public. tableau.com I'm going to leave the link in the description and from there we're going
to find the menu create and then we're going to click on that then we have download tblo desktop public edition so
let's click on that and then we're going to go to the middle and click on download Tableau public and now before
the download starts we have to fill out this registration Forum this is not for creating public account it's just
something before download starts so we're going to give the first name last name email and Country and then we're
going to click download the app and then the download going to start it's just 500 megab so it should not take long
time and now we have the download is done so let's click on the execution file to start the installation process
okay so at the start of the installation we are at the welcome page and here as usual we have to read and accept the
terms so you have to do that and here we have second box you can click on it if you don't want to send the product usage
data to Tableau team it's like cookies I don't mind I'm just going to leave it so we click now install and once you do
that the installation going to start it should not take long time okay so now the installation is done and Tableau
going to be launched automatically okay so let's go back to the website public. tau.com and on the
right side at the top we're going to click on sign in and then we have to click on this join now for free now we
have to fill out this registration Forum in order to create a new tblo public account so we have to enter the name the
email the password and the country and then we have to read an agree on the terms and let's click here I am not a
robot and at the end we're going to click on create my account and now we got the message to verify our account so
that means we have to check our emails in order to activate our account so let's do that okay so now after checking
I got an email from Tableau so I'm going to click on it and then I'm going to click on verify now in order to activate
our account so I'm going to click on that and then it going to send me to my account and with that we have brand new
active Tableau public account well it's like any other social media account you can add your personal informations for
example we can add our photo or Avatar so let me check what I can do over here so I have this photo from studard
television Tower it's amazing there and then I'm going to click save and we can add many other stuff so let's click on
edit profile and as you can see over here you can link your social media accounts or add your websites and so on
so let's click save if you want to learn any new tool like tblo powerbi or any other
programming languages you need always a good data set for training and practicing I start searching for good
training data sets and after a lot of research I downloaded like many many data sets but I was not happy with them
I didn't like them because they don't cover all the scenarios that we need for training let me tell you why this is an
issue in real projects your data going to be stored typically in data warehouses or data leaks inside many
many different tables and the first step in any visualization tools like Tableau or powerp is to connect those tables and
combine them in one big data model so training with only one table not going to help you and prepare you for real
projects and that's why I decided to make my own data sets to cover all the training scenarios and to have multiple
tables in order to learn how to combine the them in one data model and of course you can use my data set in order to
learn anything else like SQL Python powerbi and so on so let's see what I have prepared for
you all right the first thing that we going to go to the link in the description and then you're going to
land in my website where I've have collected all the course downloads and materials in one page so for example
you're going to go and download the training data sets we have here some important links the three sheet sheets
and many Skitch notes that I have prepared for this course and the as well you're going to find for each section
what are the important links and sketches and as well the Tableau files this link going to be available for you
after the course as well so you can always come back here and download the stuff that you need and of course for
free but now what we're going to do we're going to go and download the training data sets that we need for our
course and here as you can see we have two zip files one for the non EU and one for the EU so if you are currently in
Europe what you're going to do you're going to go and download these data sets but for all other countries you're going
to go and download the first data sets the non-e training data sets and now you might ask what is the differences
between them well it's about the decimal numbers since in our data set we have different decimal numbers like the sales
in different countries we have different representations of the decimal numbers so all the European countries they use
for example the comma to separate the decimal from the whole number but in many other countries USA in Asia we have
the dot in order to separate the decimal number from the whole number and if you are using the wrong format what's going
to happen table will not understand understands that this field is a decim number and it going to convert it to
string so now dep bend on your location go and download the data sets for me I'm in Germany so I'm going to go with the
second one and as I said it's depend on your location so let's go and click on that so next what I'm going to do I'm
going to go and grab the zip file and put it somewhere s so I don't want to leave it underneath the downloads so I'm
just going to create a safe path for that and then start extracting the data okay so now let's go and unzip the files
so I'm going to go and extract all of them okay so so now let's go inside it and check the data so here we have three
different data sets the first data set the table project sales dashboards we're going to use it in the last section once
we start building our projects then we have two other data sets the big data sets and the small data sets we're going
to use these two data sets in the whole course so the small data source and the big data source they are very similar so
now you might ask me why do we have two data sets Okay so now let's open both of them and see what do we have inside them
so as you can see we have almost the same tables so customers we have orders products and so on so they are almost
identical and now you might ask me why do we have two data sets well because we have many different types of
calculations and functions for example some calculations going to change the data at the role level and it's better
to have a small data set in order to understand their results easily and in the other hand we have calculations like
aggregations on the table LOD it's better to have many data in order to understand how it works and that's why I
have decided to have two data sets in order to cover all those scenarios and another thing about the data sets is
that the file type is CSV we have only one Json over here so you can use either Tableau public or tblo desktop in order
to follow me in the course all right so now I'm going to walk you through the data model of our
data sets here we have three typical tables our data sets contain information about the superstore use case it is
simply sales transactions of customers ordering products by a company it's classic and very easy to understand the
first table in our data model is the customers table it contains all customer informations such as the name of the
customers their locations and their score in the small data set we have five customers and in the big one we have
around 800 customers and the second table in our data model is the orders it contains all the orders placed by the
customers so we have informations like the order date sales quantity and profits in the small data sets we have
10 orders and in the big data set we have a around 5 years of data and that's really helpful once we start building
clusters and the third table in our data model is the products it contains all the products that we find inside our
super store so we have informations like the product name category and the subcategory in the small data set we
have only five products in the category Monitor and accessories but in the big data sets we have more than 2,000
products with categories and subcategories all right so now we have those three tables but as well we have
relationships between them like for example example there is a relationship between the orders and customers they
can be connected using the customer ID and if you check the orders and products you can find another relationship
between them where you can find the product IDs in both tables and with that we can make a relationship between the
orders and products all right guys so I leftt all those informations in my website you can find there all the links
to the data sets that I found during my research so you can go there and check them if you want
[Music] okay everyone so let's start Tableau public desktop if you don't have it open
already and then in the starting page we're going to go to the left menu to connect TBL our data so click on text
file and now we're going to go and find our file the customer CSV that we just downloaded and now we can see the
customers data inside Tableau so let's move to the worksheets I'm going to click on the orange tab over here sheet
one to create a new worksheet and now we're going to build our visualization in Tableau we have only to drag and drop
so from the left side let's drag and drop the country in the columns and let's get another one let's move the
count to the rows all right so that was it we have our first visz and here you can see in this visual how many
customers we have in each country so with that we are done building the workbook and now it's time to share it
so sadly in tblo public we cannot download it locally at our PC but I'm going to show you work around later so
now the only option that we have is to publish it to our new tblo public account okay so now in order to do that
let's go to to the main menu over here then click on files and then we're going to click on save to Tableau public for
the first time you have to sign in with Tableau public account that we just created all right so now let's click on
sign in and now we have to give it a name and I call it my first viz and once you click save Tableau public desktop
going to start publishing our workbook to Tableau public and once it's done with the publishing a web page can open
automatically directly showing your viz in your public account so here is our viz let's go back now to our home page
and as you can see over here we have our first Vis published to tblo public and let's go inside it again and now
everyone in the world can see your viz interact with it and even download it so let's see how we can download that there
is download icon over here then click on that and now you can select the file format that you want let's select the
last one is Tableau workbook so click on that and then click download and now we will get the Tableau file twbx where we
have our data and our visualizations inside it so if you open it you can see our work again and this is the workr
that we can use in order to save our work locally at our BC in tblo bu now I remember 2014 the first time I
open Tableau I was overwhelmed with all icons and parts that we have in Tableau interface and navigating through Tableau
Pages was very confusing for me at the start and that's why I'm going to take you in short tour in Tableau interface
so let's go okay so now let's go and start Tableau and now the first thing that I want to
show you is that the whole thing the whole file we call it a workbook and the workbook is like any other book it
contains different sheets and the Tableau workbook contain three main Pages we have the start page it is the
main page where you can connect our data to Tableau and then we have the data source page it is the place where you
can connect and combine your tables together and do changes to the metadata like green naming columns and so on and
the third page where you're going to spend most of the time is the workspace page it is the place where you're going
to build your data visualizations all right so now we're going to learn how to navigate through those pages and how to
switch between them okay so once you start Tau you will be in the welcome page the start page
and now if you want to go to data source page we have to connect something so let's go again to the left side over
here connect to text file and then select our file customers and open once we do that we're going to land
automatically in the data source page and now if you want to go back to the start page so in order to do that we're
going to go to this Tableau icon over here on the left side so if we click on that we're going to go back to the start
page and if you want to go back to the data source page we're going to click on the same icon so click on that again and
we are back to the data source page so with this icon we can always go back to the start page of Tableau all right so
now let's see how we can go to the work work space page in order to do that we're going to go to the bottom over
here you will find different tabs the first one is always the data source tab this is exactly where we are now at the
data source but now if we select the sheets tblo going to take us to the workspace page and if you want to go
back to the data source page there is two ways to do that first we can stay at the bottom over here and we can select
the data source tab so by clicking on that we go back to the data source and the second option is that add a data
pane so if you go to the left side over here you can see our data source customers and if you double click on it
we're going to go back to the data source page okay guys so that's was it this is how you can navigate through
Tableau Pages let's have now a quick overview of each page okay so let's start with the first
page the start page we can see here three panes connect open and discover in Connect we can find all different types
of data connectors and in tblo public we have around 10 that's enough for the training but in tblo desktop we have
over 90 data connectors and now in in the middle we have open once you start Tableau for the first time this section
going to be empty but as you start creating new workbooks Tableau going to start showing you the most recently
opened workbook and this is really nice to have quick access to our workbooks here we have only one the first pH that
we published before and in the right side you will find this cover you will find different stuff from Tableau team
like blogs news training tutorials and so on and now in the bottom you can see informations about Tableau Software for
example now it shows that we can upgrade to Tableau desktop or later once Tableau releases new version of Tableau you will
find information here to update your Tableau but since we just installed the most recent version of Tableau it
doesn't show it okay so that was it for the start page let's jump now to the next one we have the data source page
and by now you should know how to go there by clicking on Tableau icon okay so what do we have here in the
data source page on the left side you can find all informations about our data in connections you can find the
connection informations and files you can find all tables that are inside our data and then in the middle we have the
data source name and then over here we have the area where we're going to build our data model and it contains two
layers The Logical layer and the physical layer I'm going to explain that in the next tutorials don't worry about
that and beneath that we have the data Grid it's going to show us a sample of our data and as default it going to show
the first 1,000 rows of data and in the left side we have another grid this is the metadata grid it show us more
details about the tables Fields all right so that's all for now we're going to move now to the next page the
workspace page and we can do that by selecting the sheet tab okay so in the workspace page we're
going to spend most of our time here building our visualizations that's why we have a lot of icons and stuff around
so let me quickly guide you here in this interface okay so we're going to start on the top we have the toolbar it
contains a lot of icons and those icons are the most frequently used functions in tableau so as you are building your
visualizations you have a quick access to those functions and as you might already notice there's some functions
that are not selectable well you have to understand here that in Tableau if something is grayed out that doesn't
mean that this feature is not available in Tableau public but it means it is not relevant for the visual now so for
example if I go over here it's going to sort the visual and since I don't have anything so it's not relevant to sort it
let's check the other icons we have the Tableau icon it's going to take us to the start page you know that already we
have the undo and redo the last action in the visual and as you can see as I'm hovering on the icon Tableau going to
give me short description of the function so here we can create a new data source or over here we can create a
new worksheet and so on so just hover all the icons and you will see the function all right so now let's move to
the left side we have here two panes the data Pane and analytics pane as default table going to show us the data pane but
if you want to go to the analytics pane just simply click on it so you can switch between them by just selecting
them so let's see what do we have here in the data pane the first thing is the data source that contains our data and
below that we can find the tables inside this data source we have currently only one table the customers and we can see
over here the fields or columns inside our tables and here we have as well a search field sometimes our data source
gets really big and we're going to have a lot of fields so this is really nice way to search for specific field okay so
now let's go to the analytics Spain and you can find over here predefined functions that you can add to your
visual like adding a an average line or doing clustering or even you can create your own reference line really nice
stuff okay so now I'm going to switch back to the data Paine all right so now let's move to the middle and you can
find over here different shelves and cards we're going to use them in order to build our visualizations and
everything works here with drag and drop so let's start with the first one the rows and column shelves the visuals of
Tableau they have two Dimensions the rows and columns like any other tables so if you put fields in the column shelf
it's going to create a color of the table while if you put fields in the row shelves it's going to create a row of
the table easy stuff so now let's have an example okay so let's go to the left side and we're going to drag and drop
the countes on the columns and with that we Define The Columns of the visual over here so now we're going to have
something in the rows let's take the counts and drag and drop it on the rows and with that we Define the visuals
columns and rows so if you want to S between them you can go to the toolbars over here and click on this icon and you
can switch between them very easily if you have a lot of columns I'm going to switch back and now we can add more
columns and more rows so for example let's take the city drag and drop it on the columns over here so you can have
multiple stuff and now if you want to remove one of those columns you can do that by drag and drop on the empty space
okay so let's move to the pages shelf you can use it to split the current visual into series of pages if you want
to analyze something like step by step and take it slowly so let's have an example okay so let's take again the
customer count drag and drop it on the pages and now as you can see on the right side we have a new window to
control the pages and now we are at the first page where we have countries with only one customer so if we click over
here on the right side you will get the countries with two customers and so on and now for the next example I'm going
to remove it so I'm just going to drag and drop in the empty space all right so let's move to the next shelf we have the
filters you can use it in order to filter our visual for example let's take the countries drag and drop it in the
filters and now you can here decide which country is going to stay and which country going to leave the visual so now
if I select for example let's remove friendss and click apply you can see our visual don't contain now the country
friendss and now I'm going to remove it again from the shelf by drag and drop in the empty space and then we have the
marks card you can use it in order to design the visual so for example we can add new colors so if we drag and drop
the countries on top of the colors we will get the color for each country or we can change the size of the pars
either make it small or big or we can add labels and so on okay so now let's move to the middle of course here we
have our view it contains visualizations or we call it vises so first we have the title and you can change it by double
click on it let's give it the name for example customers by country and then click okay okay and
below that we have our visualization and it contains different stuff for example we have the headers and here we have the
countries and as well we have the axes now the intersection between those fields are the marks
and those marks could be like bars in this example or could be a line or circles or any other shape and now if we
check the bottom of Tableau interface you can find status bar it contains a lot of details about our visual for
example it says we have three marks of course we have three pars and we have one row and three columns and the total
number of customers is five and now let's add more stuff to the visual to see how those status change so let's
take the scores drag and drop it in the rows and you can see here we have now six marks we have six bars we have two
rows and three columns and those statos are really important once your visualizations get complicated so now we
have very simple one we can count it and see we have six parts but if we have a lot of Dots and a lot of points it's
really hard to count them so it's really nice to check the status bar to see details about our visual all right so
now let's move to the right side and we're going to go to the show me icon so select that now you will get different
visualizations that offers and by just clicking on them you're going to switch the whole visualizations in our view so
here we can switch it to tables or to P chart or to tree maps and so on so now just go and explore those different
visualizations and you might already notice that some of them are grade outs we cannot use it here again it's
available but we don't have the requirements to use it so for example if you go to the line chart here Tableau
tells you what are the requirements or what Tableau needs in order to build this visualization so it needs one date
it doesn't need any dimensions and it need at least one measure and currently in our view table cannot create it
because we don't have any date field in our view all right everyone so that was the main component of the worksheets now
before we go to the dashboard I'm going to do a few stuff you can follow me okay so I'm going to undo those
visualizations and go back to the bar and then I'm going to create a new sheet so I'm going to click over here create a
new worksheet and then I'm going to take the countries and this time I'm going to take the scores
over here and then I'm going to use the byy charts and over here I'm going to put
some labels on it okay so that's enough let's go now to the dashboard we can do that by creating new dashboard on the
icon over here and now we are at the interface of the dashboard I'm not going to explain
everything over here it's just important to understand that in the dashboard we can start combining different sheets in
one place so we can drag and drop the sheet number one where we have the customers by country and then we can
take the sheet number two just place it somewhere over here and then I have in one place two visuals the sheet number
one and sheet number two and this is the main job of the dashboard all right everyone so now I'm going to show you
the last type of sheets we have the story in order to create a new one we're going to go to the bottom over here and
click on this icon and with that we have created a new story and stories in Tableau they are like sequence of
visuals and and we use it usually for presentations if you want to tell a story from our data all right so what do
we have over here in the left side we have the visuals that we created we can see the worksheets and as well the
dashboard and then over here we can add a new story points and in the middle we have in this section like navigator to
go through our story and then here we're going to present the story or the views so what we're going to do now in the
first one we can to drag and drop the dashboard let's do that and now we're going to add a Next Step by adding blank
over here and then we're g to take the sheet number one and then we're going to add a new one blank and then sheet
number two so now we have like story it starts with the big picture with the dashboard and as we go through the story
step by step we go more in details in each visual it's really nice way to present or to tell a story using our
visuals all right so now we have the Tableau Software installed we have the two training data sets the public
account to share your work and everything is ready to start learning Tableau so with that we have finished
this section where we have prepared your environment to practice Tableau and in the next section we will do deep dive in
the Tableau Data source to learn how to build a data model in Tableau by combining
tables data modeling in Tableau each successful dashboard or charts in Tableau going to be based on a solid
data model and having data modeling skills is essential for each Tableau project or business intelligence
projects so that's why we're going to start learning the fundamentals of data modeling including the star schema and
the snow Fleck schema and then I'm going to introduce you to the Tableau Data modeling where you're going to learn the
physical and The Logical layers and then we're going to learn the different methods on how to combine tables in data
modeling using joints Union relationships and data blending and of course in order to understand the
differences between them we're going to compare them side by side and of course I'm going to guide you in when to use
which methods and at the end we're going to go and build two data sources based on our training data sets so let's start
with the first topic where we're going to understand the fundamentals of data modeling so now let's let's
go in real projects your data going to be sted typically in data warehouses or data Lakes inside many many different
tables and the first step in any visualization tools like Tableau or Barbi is to connect those tables and
combine them in one big data model so let's start with the question what is data modeling data modelling is
the process of organizing and representing data in a clear and understandable way each data model has
entities entities could be things like customers and products or events like orders and inside those entities we have
informations and we call them attributes like the first name and the last name inside the entity customers and we
describe in the data model how those entities are connected or related to each others and we call it relationships
this data model this visual representation of the data makes it easier for us and for programs to
understand the data which is really important for making decisions and improving performance of the
business all right so we have three different types of data models at different levels of abstraction first we
have the conceptual data model this type is high level representation of the data model without going in details on how
the data model is implemented it's like a map that shows the important entities and the relationships and we usually use
this type to explain the data models to business analysts and stockholders to understand the big picture of the data
the second type is The Logical data model in this data model we go more in details on how the data is structured
and organized we Define in this model the attributes of each entity and it includes as well constraints and more
details about the relationships between the entities this data model is usually used by database designers and
developers as a blueprint for the implementations and the third type is the physical data model this type
represents the actual implementations of the data model it includes all the technical details about how to store the
data like the data types of the attributes the primary and foreign Keys indexes and so on this data model is
used by developers to create and manage the databases all right so let's summarize the conceptual data model
shows the big picture of the data The Logical data model provide a blueprint for the implementation
and the physical data model shows how the data is implemented in the databases and Tableau did adopt both the logical
and physical data models in the data sources but we don't have conceptual data model in Tableau don't worry about
it I will show you more details later all right so now for analytics and especially for data warehousing and
business intelligence we need special data models that are optimized for queries and for analytics it should be
flexible and easy to understand and for that we have two special data models first one is the star schema star
schema has a central fact table and surrounded by dimensional tables the fact tables contains events and the
dimensions hold descriptive information the relationship between the fact and the dimension tables form a star shape
and that's why we call it a star schema and the other data model we call it snowflake schema it is very similar to
Star schema but the dimensions here are breaking down into subdimension normalized tables or Dimensions means
that those tables are broken down into small pieces to avoid having big tables or big Dimensions which leads to many
data duplications and slow performance the shape of these data models looks like
snowflake so star schema is a simple and easy to understand data model and we usually use it if our data set is small
or medium in the other hand the snowflake schema is more complex but it eliminates the duplicates and reduces
the storage spaces and we usually use it if we have large data sets all right so the data sets that I've prepared for
this Tableau course are using the star schema data model just to keep it simple and easy to
follow all right so our data model has a name and we call it star schema if you're going to work on real projects
you're going to hear about the star schema a lot so star schema has mainly two types of tables facts and dimensions
for example we have the table customers it Des describes each customers by their first name last name country and so on
so customers is a dimension table and we have another dimension table in our data model it is the products so products
table describes as well each product by their name and category so it is as well a dimension all right so now let's talk
about the second type of tables in the star schema we have the facts for example let's have a look at the big
table in the middle we can see three things you can see first a lot of keys to the other dimensions we have the
order ID customer ID product ID and we can see dates so we have the order date the shipping date and the third thing we
can see a lot of numbers so we have sales quantities profits we call them as well measures so if you see those three
things that means we have an event or fact table so facts Connect dimensions together it has dates and as well
measures okay so to summarize how do we decide if a table is dimension or fact if you have a table that contains
informations about a physical person or an object like employee customers product s then this table is a dimension
and usually they are small tables and in the other hand if you have a table that contains events for example we have
sales orders logs ATM transactions so any tabl that has events transactions and has time in it we call it facts and
usually they are really huge tables okay so in our data model in the data sets we have two Dimensions we have the
customers and products and in the middle we have our fact the orders all right so now if you hear in your project someone
talking about star schemas and so on you know exactly what they mean it's very important Concept in analytics and bi
world if you are using Tableau or power Pi okay so once we connect our data to Tableau we have to create a data model
in our data source and if your data contains only one table then your data model is very simple you have single
table in your data model but in real life projects things get more complicated where you have multiple
tables and Tableau here offers four different methods of how to combine and connect your tables we have
relationships joins Union and data blending and now before we start doing deep dive in those four methods let's
first understand the data modeling in Tableau in Tableau Data model we have two layers we have the physical layer
and on top of it we have the logical layer in the physical layer we might have some couple of physical tables and
we can combine them in Tableau using two methods either joint joining the tables or using Union between them and now
let's move to The Logical layer it is the top level layer and provide us like an abstract to hide all the details in
the physical layer this is especially nice if we have a lot of tables in the physical layer so once we are building
our visualizations we don't want to see all those tables in the physical layer so The Logical layer going to provide us
like an abstract or going to hide all those details so the result of merging the tables using join and Union in the
physical layer going to be presented in The Logical layer with single table flat table and we call it a logical table so
that means we're going to have two logical tables the first one going to present three tables after doing the
join and the second one going to present two tables using the union but we still have in data modeling to connect those
two logical tables and in Tableau we have only one method to do that and we call it relationships and it's very
important to understand that in The Logical layer we cannot merge tables in one table so after reconnecting them
using their relationship between the two logical tables the table is going to stay as it is and nothing going to be
merged we just describe the relationship between the two logical tables and now back to those two layers both of the
physical layer and The Logical layer we can find it inside Tableau Data source and as you know on top of the data
source we have our visualizations and you can see in this example only the tables from The Logical layer and you
can start building your visualizations using the data available from The Logical layer but sometimes as you are
working with the projects you build another data source with another data model and here in this example it's
important to understand that not all logical tables comes from the physical tables they could come directly from
your Source system and now in order to build one visualizations from both of the data models and the data sources we
have somehow to connect those two data models or data sources and we can do that in the visualization level where
Tableau offer us the last and very unique method of connecting and combining tables something called Data
blending so by looking at this you can see that will offer us four different methods of how to combine and connect
tables in different layers and different levels so in the physical layer we have the joints and unions we have in the
logical layer the relationships and at the visualization level we have data blending all right so now let's see in
Tableau how we can navigate through the physical and The Logical layer we are currently at the data source page and as
a default we're going to be add The Logical layer in the data model so that means anything that we drag and drop in
our data model going to be considered as a logical table so the customers is a logical table let's take another one
let's take the orders drag and drop it over here so this is our second logical table and as you can see Tableau did
create between them a relationship because at the logical layer we can do only relationships so now we are at The
Logical layer how we can go to the physical layer in order to do that we're going to go inside a logical table so
let's go to the customers and double click on it once we do that we're going to go to the second layer we are are
inside the physical layer now so table going to tell you over here the customers is made of one table because
we have only one physical table so now anything that we drag and drop in the data model going to be considered as a
physical table so for example we can take the customer details let's drag and drop it over here and by default Tableau
going to create between them not relationship it going to create a joint between those two physical tables and of
course we can do a union between them so in the physical layer we can do joins and unions and as you can read over here
it says the customers The Logical table customers is made of two physical tables and if you have her on this icon you
will see exactly that so we have two physical tables defines the logical table customers and now if you want to
go up back to the logical layer we can do that by just closing the physical layer so let's click on that and now you
can see that the customers has a new icon it says in the physical layer there is like join and we get more
informations if we hover on the tables it says logical table customers that is made of two physical tables the
customers and the customers details so that means the data in The Logical tables comes from the physical layer but
if we go to the orders over here you will see no physical tables the data comes directly from the original tables
and with that we have learned how to navigate through the physical and The Logical all right so let's start talking
about joining tables we usually have two tables table a and table B and if you want to combine them in one big table
then then we can use join between them the first thing to understand is that once we use joint between two tables
then we have two sides table a going to be the LIF table and table B going to be the right table so now what going to
happen after we join the tables all the fields from the left table will be at the output and then all the fields from
the right table will be added next to it so joins combines the fields or The Columns of two tables so now in order to
do joints we need two things first we need the key field it is a field that you can find it in both St tabls and
after that we have to define the type of joint and we have to choose between four different types of joints we have the
inner join the left join right join and full join and if you know SQL then you know those types it's exactly the same
logic but let's have a quick examples to understand the four types of joints all right so now we have this
example where we have two simple tables we have the customers names and the customer's age and we want to combine
them in one table because it makes no sense to have two tables about the customers so we want to make one
customer table and we want to combine them in the first table we have the ID and the names and the second table we
have as well the IDS and the age so it's really easy the key for this joint is the customer ID now let's see the
different output using those different types of joints so let's start with the first type of joint the inner joint
inner joint says the output going to show only the matching rows from the left and from the right so that means
any unmatching rows will not be presented at the output so let's see how this works so the first thing that's
going to happen is that we're going to combine first the fields so first we're going to start with the left
side and then the right side and now we're going to start matching the rows we're going to start from the left side
do we have the user ID one in the right side as well so we have a match so in both tables we have the customer ID one
so this we're going to see it at the output and then we proceed on the left side do we have customer ID number two
as well on the right side you see we don't have it we have only the customer number three that means two is not
matching on the right side and as well the customer three is not matching on the left side so that was it if you use
inner join in this example you will get only the customer ID number one since we find it in both tables okay so let's go
to the next one we have the left join left join says we going to have everything from the left table without
checking anything but from the right table we're going to have only the matching rows so if we do left joint
between those two tables we're going to have the following output so first we're going to have the fields from the left
table and the fields from the right table near each other and then we going to have all the customers from the left
table without checking anything so everything going to be presented over here those two customers and then from
the right side we going to have only the matching rows so that means do we have the customer ID number one on the right
table yes we have it then we're going to have it at the output but the customer ID number two we don't have it at the
right table which means it's going to be empty and empty means nulls so here we going to have the values of nulls in
both of the field ID and as well in the age and that's it this is the output of left join all right so now we're going
to move to the next one we have the right join you might already understand how it works so we're going to have all
the RADS from the right table and only the matching draws from the left table so let's see how the output going to be
if we do right joint between those two tables as usual we're going to have all the fields from the left all the fields
from the right and we're going to have all the rows from the right table without checking anything so we're going
to have those two customers and then we start matching from the left side so do we have the customer number one yes we
have it so we're going to add it over here do we have the customer number three so as you can see we have only the
two that means we don't have informations and we going to have the nulls so those going to be empty
and that's it so it is exactly the opposite of the LIF join and now to the final type of join we have the full join
full join means everything from left and everything from right without missing anything so let's see what going to
happen if we have full joint between those two tables so as usual we start with the fields so from the left and
from the right and then we take everything from the left side so we take those two customers over here and from
the right side we're going to have the matching Row for those two customers so so for the ID number one we have this
one but for the two we don't have any matching rows so we're going to have nulls over
here but as you see we don't have everything from the right side so the customer ID number three is missing so
that's why using full join we're going to have those informations over here and then we're going to match it as well
from the left side so do we have any customer number three on the left side we don't have so that means we're going
to have NS as well so now by checking the output you can see we have everything all the data from left all
the data from right and where there is no match we're gonna have nulls so as you can see you need to be really
careful with the type of joint you are using because using the wrong one this could cause of losing data and if you
want to be safe and you don't want to lose any data then you have to use the full joint but sadly full joints are
very slow and you're going to end up having very big tables especially if both tables have a lot of unmatching
rows and now I want you to understand how joints Works in Tableau and what can happen in the background Once We join
tables so we have the data source we have the visualizations and inside the data source we have the physical layer
and The Logical layer in the physical layer we're going to join both of the tables A and B and once we do that tblo
going to create one new combined table A and B in The Logical layer this table we call it a logical table which contains
data from both tables and then in the visualization layer let's say we want to select select the fields of F2 and F4 so
Tableau going to query the data source and the data source going to get the data from the new combined logical table
ab and then send the data back to the visualizations so as you can see the interaction between the visualizations
and the data source going to be at The Logical layer so the physical layer going to be completely out of the
picture and that's simply how joints Works in Tableau all right so now how we can do
joints in Tableau let's say that we want to join the table customers with the orders so first we're going to go to the
left side over here drag and drop the customers and the joint is going to be done at the physical layer so we have to
go there so let's go inside the customers and now we are at the physical layer and we're going to take the orders
and just drag and drop it over here at the empty space and with that Tableau as default going to create an inner join
between the customers and the orders and if we want to customize the join we going to go over here at the icon and
click on it and we have here two things to do first we're going to define the type of join as we learned we have the
inner left right and full outter join you can just click between them and see which data going to be missing and which
data can to be presented as the example that I showed you so I'm going to stay with the inner join and the next thing
is that we're going to define the key for the join so tblo did understood there is customer ID from the left there
is customer ID on the right and this is the perfect match which is correct but let's say it was wrong and you want to
choose the correct key for the join what you're going to do you're going to go to the left side over here click on the
Arrow you will get all the fields from the left table and select the correct one in this example the customer ID is
correct so I'm going to stay with it and you'll go to the right side you have as well the same icon over here and you
will get all the fields from the right table and you select the one that suits you and one more thing your key for the
joint could be not only one field it could be multiple Fields so you can add more Fields over here so you go to the
next row and select the next field for the join but in this example we have only one key so I'm going to close this
we have set up the joints we going to stay with the inner join and we can go back to the logical data model and as
you can see the table over here has the icon of join it tell us that this logical tables is a result of joining
two tables and that's it this is how you can do joins in Tableau all right so that's all for joints next we will learn
the second methods how to combine tables using Union [Music]
[Music] all right so now let's talk about Union let's say that we have two tables and
both of them has exactly the same columns sometimes it makes sense to combine them in one big table and we can
do that using the union so once we do Union what's going to happen The Columns and the rows of the left table going to
be presented at the output and from the right table only the row is going to be append at the output beneath the first
one so Union going to combine the rows of two tables and in order to do the union correctly we have two requirements
first both of the tables should have exactly the same number of fields and second the field should have exactly the
same data types so as you can see we don't need a key between those two tables it's not like the
join all right so now let's have a quick and very simple example about the union we have here very simple two tables the
orders of 2022 the orders of 2023 and as you can see both of the tables has exactly the same structure so we have
two columns the ID and date in both tables and it makes sense to merge them in one table we call it orders so if we
do Union between them what can to happen at the output it's going to start from the left table and it going to take the
fields first so the ID and date and then it's going to take all the rows from the left side and put it at the results and
now from the right table we will not take again the fields because we have it already from the lift table it's going
to take only the rows and abundant at the end of the table so it's going to take the two orders three and four and
just put it beneath the table over here and that's it it's very simple and easy it just need exactly the same number of
columns or fields and exactly the same data types all right so now let's understand
how Union Works in Tableau and what's going to happen in the background once we do Union so we have here again our
layers and Union is very similar to join in the physical layer we have our tables A and B and once we do Union between
them Tableau going to create a new combined logical table where it going to combin the rows of both tables and then
in the visualization level let's say that we take the field F1 tblo going to send a query to the data source and data
source going to ask The Logical table to get the data and once Tableau get the data from the data source it's going to
present it as the visualization and as you see again here the interaction is between the visualizations and The
Logical layer all right so now let's see how we can do Union in Tableau we're going to work
with the two tables orders and orders archieves both of them has exactly the same number of fails and as well exactly
the same data types so in order to do that we're going to take the orders drag and drop it on the logical layer but you
know we can do Union only in the physical layer so we have to go inside the orders double click on it and now we
are at the physical layer let's take the second table the orders archieve but now instead of dropping it at the white
space because tblo then going to create a joint we don't want to do that we want to create a union just drag and drop it
beneath the table and as you can see table going to say drag table to do Union so if we just place it beneath it
tblo going to do Union between those two tables and as you can see there is two lines gray lines indicates that there is
Union and if you want to check that you can check at the result over here the data we will get a new field called
table name and you see some records comes from the orders and other records comes from from the orders archieves
which indicates that we have one combined table of both of the orders and the orders archieve let's go back to the
logical layer so I'm going to press here the X and as you can see we have a new icon over here it indicates that we have
a union and as you can see the tool tipe of Tableau it explains everything so we have a logical table called orders it is
the result of Union table orders and orders archieve this is one way of doing Union between two tables in Tableau
there is another way to do that so let me show you how to do it first I'm just going to move it drag and drop it
somewhere over here and as you can see on the left side we have something called New Union so double click on it
and you can see we have here two options the manual and as well the automatic the manual we're going to get the result
exactly like we just did so what we can do we can just drag and drop the tables over here the orders and the orders
archieve and then click okay with that we get exactly the same results without going to the physical layer and drag and
drop two tables and put it exactly underneath the table so this is nice way to do Union between two tables you can
check that by just going to the physical layer so double click on it and as you can see we got exactly the same results
and here we can check the table name we have orders and we have the orders archieved all right so now let's check
the second option where we can do Union automatically I will go back to the logical layer and just remove the union
over here let's start a new one from the scratch and now we're going to go to the automatic so what do we have over here
imagine that we have around 100 tables about the orders and this is very if you are not working with databases you are
working with files and the files has limitations so what we're going to do we're going to go and split the files
after day after months after year and so on so we end up having a lot of files and it is very painful if we're going to
go and drag and drop all those files in Tableau to do Union and instead of that we're going to Define for Tableau or
Rule and Tableau going to go and search for all files that follow the rule and do Union between them so what that means
for example we have here two tables the orders and the orders are she what is the naming convention over here it both
of them starts with the orders so I could have like a third table called orders uncore 2022 orders uncore 2023
and so on so there is a rule I'm following here in my naming convention and I can specify that in Tableau so
let's see how we can do that so over here the first option is going to include or exclude I'm going to leave it
as include and now I'm going to specify the rule so it start exactly with orders and after this word it doesn't matter
what comes after that it could be underscore 2022 2023 or nothing and so on so anything after that doesn't matter
so what we're going to specify after that a star Stars means anything after orders and then we have some options to
tell Tableau where exactly to search either at the sub folders or at the parent folders I'm going to leave it as
it is and then click okay so now we have a union let's see what tblo going to say says we have a logical table called
Union and it says we have many Union table because we have the automatic way of doing that and now let's check
whether Tableau did that correct so as you go to the right side here in the overview you find we have a new field
called path it is the path of the files so let's see that I'm going to go to the sheet one here and just drag and drop
the pass to see just the files so as you can see tblo did it correctly we have the orders archieve and the orders it's
really nice way if you have a lot of csvs and excels to do it automatically instead of drag and drop all those
tables usually in my projects I never use this because all the data is prepared in the data warehouses or in
the data lake so with that we have learned all the different options on how we can do Union in
t all right so now let's talk about relationships in 2020 tblo introduced a new methods on how to combine and
connect tables together and they called it relationships they made it even as a default methods on how to connect tables
since it is very fast and flexible so what is relationship ships and how it works in Tableau it is
completely different than joins and Union if we have in the logical Layer Two logical tables A and B we can
connect them at this layer using the relationships think of the relationships as a contract between two tables and
when Tableau uses the data from those tables it has first to check the contract in order to understand how to
generate the queries and now it's very important to understand that once we connect the tables using relationships
the tables can stay separated from each others and Tableau will not create a new logical table so everything going to
stay as it is without any changes and here we just describe the relationships between two tables so now in the
visualization level if we take the field F1 from table a and F4 from table B what can to happen first Tableau going to
check the contract in order to understand how to generate the queries and then it going to send the query to
the first table and then it going to send another query to the table B in order to get the data for F4 and then
the data going to be combined at the visual I ization level and not the logical
level all right so now let's see how we can create relationships in Tableau it's really easy so we're going to stay at
the data source page and as well at The Logical layer we will not go to the physical layer and all what we need is
two tables so let's take the orders drag and drop it over here in the data model and then let's take the customers so now
as you can see as a moving there is like a noodle or relationships so let's drag it here and tblo going to automatically
create create relationships between the orders and the customers and now how we're going to configure and set up the
relationship so let's go to the Noodle over here and just click on it and then there will be no new window or something
for the setup we're going to go to the metadata over here if you don't see the information like this then you can go
over here and you will see like the relationships and The Logical tables so make sure you are selecting the
relationship and there is like three things that we're going to set up as a relationship first it's going to be the
key it's like the join key it is Comon filled between between the two tables so now as you can see over here from the
left table we have the customer ID and the right table we have the customer ID and tblo did automatically understand
that this field could be used as a key which is correct but if you want to change it you can go over here so we
will get a list of all fields on the left table and as well you're going to go over here you will get all the fields
from the right table and you can add more fields for the key currently it is correct so I'm going to leave it as it
is and next we're going to go to the performance options so we're going to extend the performance options over here
and we have here two things we have the cardinality and the integrity and if you leave it here as it is as a default
nothing going to go wrong you will not lose any data so you don't have to change anything here unless you want to
optimize the performance so what do we have over here we have cardinality as many or one on the left side and on the
right side you can Define the same stuff for the Integrity we have some record match and or records MKS so in order to
understand those stuff let's have an example all right so now we can to have example
for the cardinality in relationships we have two tables our orders and customers there is a relationships between them
and the key for the relationships is the customer ID and in the cardinalities there is two options either we're going
to use many or one and in order to decide which one is the correct one we have to do data profiling data profiling
means we're going to do deep Dives in the data to understand the values inside our tables and once we do data profiling
it's very easy to select whether it's many or one so now what those values means many and one there is a simple
rule for that we use many if there is duplicates in the key and we use one if the key is unique and does not have any
duplicate inside it so now let's check the example in order to determine whether it is many or one so let's go to
the orders over here and the customer ID you see in those values there is delates we have the customer ID once here and
once here as well and the customer ID two is twice so those values are not unique and contains duplicates that's
why we call it a manyu let's go to the customers over here you can see we have the customer 1 2 3 and that's it so
those values are unique and there is no duplicates inside it we don't have the customer ID one again in the table so
that means we can specify here a one so now let's go through all scenarios in order to understand what can to happen
in Tableau once you configure this all right so now let's run the first scenario where tblo going to Define it
as a default many to many relationship so we have at the left side M and on the right side we have as well
money and let's say in the visualization level we talk the customer IDs from the order and the sum of all sales and then
the name of the customer all right so now let's see how Tableau going to work Tableau first going to check the
relationships it's going to say okay it's many to many it's better to check the whole tables on the left and on the
right so we're going to start on the left side we have the customer one it's going to take it over here and it's
going to sum all the sales so since it's many table going to understand I have to check the whole table so tblo going to
scan the whole table one by one it's going to say okay we have the sales 50 the next one is not the customer one and
then go to the next it's going to skip it and then we have again the customer ID number one and it's going to do the
sum between 50 and 30 that means we're going to have the value of 80 it is the sum of the two sales and now we're going
to go to the right side to find the name of the customers it's going to check okay it is many so it's going to scan
the whole table for the customer ID one so now the first three Cod it finds okay we have the customer ID one it's going
to take Maria over here but now tblo will not stop it's going to scan the whole table since in the relationships
it's many but it doesn't make sense because the customer ID here is unique so tblo going to check whether there is
customer ID one over here and then go to the next and then it didn't find anything so it going to stay like this
and now tblo going to proceed with the next customer we have the customer ID number two we're going to have it at the
output and then we're going to have the sum of all sales so tblo going to scan the whole orders in order to do the sum
so we have over here the 20 and then we have here 10 so the sum of that is 30 TBL going to have at the output 30 so
that's it for the left table we're going to go to the right table T going to scan the recorde one by one so the first one
is not the customer ID number two we have here a match so John going to be at the output tblo going to scan the whole
table so it's going to go for the three and so on and as you can see the output is correct using the default methods of
many to many but we have a problem with that on the right table tblo is doing a full scan so with that we are losing
performance on the right side so it's better to optimize it where we going to tell Tableau if you find a customer then
that's it you don't have to scan the whole table because we have at the maximum one record of each customers
there is no duplicates and it is unique and now we have to tell somehow this information for Tableau in order to do
that we can do it in the cardinality so in the left side it's going to stay as many but on the right side we're going
to say it is one and with that t going to understand okay it is unique we don't have to scan the whole table and we're
going to win a lot of performance all right so now let's see how tblo going to work once we have it as many to one on
the left side nothing going to change because we have many so tblo going to scan the whole table so for the customer
one the result going to be the same but now on the right side things going to be changed so Tableau going to say okay
customer ID number one there is a match it's going to take Maria as the output but now Tableau going to stop Tableau
will not search for the customer ID one and scan the whole table so with that tblo will not be doing any unnecessary
Stu off and we're going to win some performance we're going to go now to the customer number two over here same
information so it t get a scan so do we have the customer number two over here no so we jump to the next one yes we
have a match we're going to take John but tblo going to stop as well and will not scan the next record so as you can
see we have exactly the same output whether you are using many to many or many to one with many to one we have won
the performance with Tableau going to stop the scan on the right side all right so now let's jump to the next
scenario where we're going to do do something wrong where we're going to say okay the customer ID on the left side is
unique and we're going to put the value of one and on the right side it doesn't matter let's have money for example so
now we are telling Tableau on the left side the customer ID is unique so you don't have to scan the whole table and
we're going to have the same example over here so let's see what going to happen on the left side tblo going to
start with the first customer say Okay customer ID one the sum of sales is now 50 because I don't have to scan the
whole table so it's going to stop at the first records and the output going to be 50 so now on the right side once we are
saying many here doesn't matter the result we going to be correct we're going to have Maria but table going to
scan the whole table so the performance going to be bad now we're going to jump to the next customer we have the
customer number two so table going to have it at the output and here again the same problem table going to say okay we
have the sale 20 the customer ID is unique we will not find it again in the same table I don't have to scan the
whole table so tblo going to take the value 20 and going to put it at the output without checking the other values
and here on the right side it doesn't matter we have John which is correct but it's going to scan the whole table so as
you can see if you make mistake here in the cardinalities you might have some problems at the output where we're going
to have some missing data and wrong informations all right so now let's run the last scenario where we have on the
left side one and on the right side as well one we're going to get exactly the same output because we have it wrong on
the left side the only good thing here is that on the right side Tableau going to stop the scan once once it find a
match so it will not scan the whole table so at the output we're going to get exactly the same informations and
here we have one to one all right so now let's quickly summarize on the left side we have two criteria the correctness and
the performance correctness is always way more important than the performance let's start with the first scenario we
have many too many relationships as you can see the output was correct but the performance was bad since tblo doing
unnecessary Full Table scan on the right side so that's why I'm going to give it okay for the correctness and not okay
for the performance for the next scenario we have many to one relationship the output was okay so it
was correct we're going to give it okay and the performance was okay since Tableau stops the scans once it find a
match so that's why we're going to win a lot of performance and we're going to give it an okay let's jump to the third
one we have one too many relationships as you can see the output was not okay it was not correct we are missing data
so we're going to give it not correct and the performance was bad because on the right side we are doing unnecessary
scans so that means it was the worst scenario over here and then the last one we have one to one relationship the
output was not correct not okay but the performance was okay since on the right side we are not doing any unnecessary
scans but to be honest correctness is way more important than the performance and that's why Tableau always recommend
to stay at many to many relationships if you are not sure because you always going to get correct answers at the
output but if your data is Big you will get some bad performance so if you want to have like good performance you have
to invest time in analyzing your data doing data profiling to understand is it many is it one and then change it but
you have to be sure about your data otherwise you will get wrong informations at your visualizations and
that's really bad so that means for this example the safe way to do it to stay at many to many relationships but the
professional one is to have many to one relationships to get good performance but this is not always a scenario just
imagine we switch the tables between customers and orders so customers is left and orders is right then one to
relationships going to be the correct one so be careful here with the sides all right everyone so now let's
understand the Integrity options in Tableau each relationship has two sides the left table and the right table when
we are changing the settings of the Integrity we limit which joints can happen in the visualization so here we
have two options some record match and or record match and with that we have four scenarios first we can choose some
record match in both left and right tables and if we do that then all types of joints are possible in the
visualization we have inner left right and full joint but now if we choose all record match on the left and some record
match on the right so what going to happen now we are limiting the types of joints to only two types inner and right
joint and the next one is going to be the opposite so we have some record match on the left and all record match
on the right what can happen again here we limit the types of joints to only two types the inner and left join and in the
last scenario if we choose all record match on both sides the left and the right then here we limit Tableau to only
one type of join the inner join so as you can see it's very similar to joints we are just defining how Tableau should
work when we use some record match we allow more types of joints and when we use the option or record match then we
are limiting Tableau with the types of joint and here it's very important to understand that we have a tradeoff if
you use or record match and go down this path you will likely experience better performance but you will increase the
risk of losing data but if you choose to use some record match and you go up you will ensure the completeness and the
flexibility but you are sacrificing some resources and performance and tblo team here decided to go with the first
scenario where you have on the left and the right some record match and I can understand that because it's more
important to have completeness and flexibility more than performance let's a look at our data so here we have
customers that didn't order anything so the customer number three didn't order anything over here and we don't have a
match of it so we can say some records matches like the one and two are matching on the left sides but some
other records does not match so we don't have an order from the customer ID number three so that means in our
database we could have customers in the customer table that didn't order anything so the correct option over here
is is some records matches now let's analyze the orders as you can see we have the customer ID number one we find
it in the customers two as well and so on so we can see that all the records or all the customers IDs in the orders has
a match from the customers well that means we can select all records match we don't have for example customer ID 4
over here which does not have a match on the right side so that means in our database all orders should comes from
our customers and we should not have any order without a known customer so after the analyzis we can say on the left
sides on the orders we have always a matching record so we're going to select all records matches but on the right
side we might have customers that didn't order anything then we can say some records matches if we do it like this we
can prevent Tableau from doing any extra stuff by analyzing the nulls like in SQL if you have full outer join you will get
like huge amount of data and sometimes if you're using inner joint or left joint or so on you will get better
performance so if you know exactly what is going on in your data then select the correct Integrity otherwise just leave
it as a default some records matches on the left and on the right you will be safe you will get correct
answers all right so now back to Tableau relationships are really easy we just have to drag those two tables and T look
going create relationships between them just get the key key between the relationships correct and everything
going to be fine and leave those stuff as a default but if you want to be like more professional and get better
performance in Tableau you have to do data profiling and then select the correct one if you are 100% sure so in
this example the orders over here has many in the customer IDs but we have on the right side one for the customers and
then for the Integrity on the orders all records matches because all orders has a customer ID in the customers table but
we might have some customers that didn't order anything so I'm going to leave it as some record matches and that's it
that is relationships in all right so now let's talk about data blending in Tableau but first some
coffee let's go all right so now let's have this example where we have in the data source table a and now in the
visualization level we want to use the data from the field F1 and you know by now tblo going to send a query to the
data Source in order to get the data of the F1 from the table to show it in the visualization and now since this data
source was the first one to be queried and to be used and Tableau going to call it a primary data source and in Tableau
anything is primary going to get the blue color that's why you will see like blue icon indicates that this data
source is a primary one and now sometimes you are in situation where we want to get the data from another data
source for example we have another data source with the table B and we want add the visualizations to show the data of
if F4 so what's going to happen tblo going to send another query to the second data source in order to get the
data of f4 and then the data can to be forward to the visualizations and here Tableau going to call this data source
as a secondary data source and it will mark it with an orange icon and now in order for this to work where we're going
to get data from two different data sources we have somehow to connect them and here exactly we're going to use the
very unique way in Tableau where we can connect data sources together using the data blending and data blending can only
be done at the visualization level on the worksheet page not in the data source so now you might ask how Tableau
is joining those tables at the visualization level well Tableau is using a LIF joint we cannot change that
sadly it is fixed since it's like a LIF joint tblo going to get all the data from the primary data source and only
the matching records from the secondary data source so now to summarize data blending is the methods of combining
data at the visualization levels from two different data sources using a LIF join and this is very unique feature in
Tableau you don't find it in any other bi tool like Microsoft powerbi you cannot for example there combine data
from two different published data sets all right so now let's see how we can do data blending in Tableau and for
this we need two data sources the first one going to be from the CSV files that we have from the small data sets so we
going to go to the text files and let's take the products over here so this is our first data source and now let's go
and create the second data source in order to do that you can go to this icon over here and then click on new data
source so let's go there it's going to be from the Json file that I prepared for you so let's go to Json and we have
the product prices so let's open that since it's Json we have to select the schema so let's go to the data over here
and click yes and then click okay so now we have two data sources in order to switch between them we go again to this
icon over here and you can see we you have now two data sources and by just selecting the data source you will
switch to it and now in order to do the data blending and to connect those two data sources we cannot do it at the data
source page we have to go to the visualization level to the worksheet page so let's do that I'm going to go to
the sheet one over here and as you can see at the data pane on the left sides we have two data sources and by just
clicking on them you can switch in order to see the tables inside them so now we have to decide which data source is the
primary and which one is the secondary for this example I will say that the product is the primary one and how we
going to do that by just using the data in the visualizations as the first data source so I'm just going to take the
product ID drag and drop it on the rows and immediately tblo going to understand okay this is the primary data source and
it's going to mark it with a blue icon over here indicating that this is our primary data source we still don't have
a secondary data source so you see there's no orange icon over here because in our view we have data only from one
data source so now in order to get the data from the second data source we we're going to switch to the product
prices and you can see Tableau immediately turn this data source as a secondary data source so you can see
over here we have the orange icon indicating that this is secondary data source and any field that we are using
it's going to mark it with orange so you can see over here the price it has an orange icon so that's it it's very
simple so now let's say that the product ID is not the key of order to join those two data sources you want to change that
in order to do that we're going to go to the data over here in the menu and then go to the edit blend relationship let's
click on that so we'll get a new window over here and here we have two options automatic and custom if you leave it as
automatic tblo going to figure out which key to join those data sources and here in this example is the product ID but if
you want to change that you can go to the custom over here it's like join you have to specify from the left and from
the right which fields are the key in order to do the join so if you want to change that just double click on it and
then you have in the left side the primary data source and the right side the secondary data source and then you
select the fields that are the key before the join so I'm going to leave it as it is and let's add another key so I
will go over here and add for example the category is from the left side and from the right side the data index which
is really wrong so let's click okay and then again okay you will see on the left side now we have another chain on the
data index and you can see it's like broken chain so that means it is not yet used in the join if you want to activate
it just click on it and you will see we have an active chain and now as you can see the result is wrong because it
doesn't make sense to use this key but I just want to show you how you can deactivate and activate the key of the
joint between two data sources by just clicking on them so now let's just correct this I want to have only the
product ID as the key for the joint so that means I'm going to deactivate the data index over here and that's it this
is how you can Define the key for the data blending and now one thing that is very important to understand is that
everything that we done in the data blending is only relevant for this worksheet so if I go to another
worksheet let's go over here and create a new one and now as you can see over here it's completely resets the two data
source we have it again but we don't have it as primary and secondary data sources that means in each worksheets we
can make a new decision so at the sheet number one the product what the primary I can change my mind here where I can
say okay the product prices now is the primary data source so if I take anything over here you can see product
prices is the primary and if I go to the product and let's say I'm going to take the product name over here products
going to be the secondary so I just switched between them depending on the requirements so if we go back to the
sheet number one we see that the product is the primary but if we go to the sheet number two the product price is now is
the primary this is really nice because it give us really flexibility where we can decide in each worksheet which one
is the primary and which one is the secondary depending on our requirements so data blending is very unique and
great way on how to connect and combine thata all right so now what is the main difference between joins and unions both
of them are very similar they're going to combine two tables in one big table but the difference here is that's how
the data going to be combined in joints the fields of both tables going to be combined so we're going to take all the
fields from the left side and beside it all the fields from the right sides so the results we're going to get one big
wild table but in the other hand in the unions two tables going to be combined but instead of combining the fields here
we're going to combine the rows of both tables so we will get all the rows from the first table and beneath it all the
rows from the right table but both of them has exactly the same columns so joints combines the fields and Union
combines the rows all right so that was the main difference between joint and Union all right so now the question is
what is the main difference between joints and data blending data blending is like a LIF joint but the main
difference here is that when the aggregation is going to be performed in joints the data going to combines first
and then the aggregation going to happen but and data blending is exactly the opposite the aggregation going to happen
first and then the data going to be combined so now let's have a simple example in order to understand what this
means okay so again we have our tables customers and orders first we're going to do the lift join and afterward we're
going to do the data lending between them in order to understand the differences between them in the output
all right so now we're going to start with the left join you know left joint all the data from the left sides and
only the matching on the right side so we start as usual by combining the fields from left the fields from right
and we start recode by recod so we're going to take the customer number one and we're going to search for the
matches we have two rows on the orders so that means Maria going to be twice in the output because there is two orders
and then we're going to go to the next one customer ID number two we have only one order for that we're going to have
it at the output and George don't have any orders so that means we're going to have nulls null here here and here so as
you can see with the lift joint first we combine the data the row data without doing any aggregations and afterward in
the visualizations we can find for example the sum of sales or the average and so on and now let's check the data
blending how it works all right so now let's say we have all the fields from the primary data source and beside it
all the fields from the secondary data source and this is like left join we're going to take all the data from the
primary data source so we're going to get all the three customers over here but the main difference here is that
there will be no duplicates as you can see we have here Maria twice but in data blending you will not get any delates
and now here comes the difference before we start getting the data from the orders from the secondary data source an
aggregation can to happen so for example with the customer ID number one we have two rows the two rows will not be
presented at the output first it's going to be like an aggregation and now it's very important to understand that the
fields in Tableau are splitted between dimensions and measures in the next tutorials I'm going to explain that in
details but now the measures can be aggregated the dimensions will not be be aggregated so for example the customer
ID it is not measure it is a dimension so Tableau cannot aggregate it but since we have it twice the same value tblo
going to write here one and then the next one we have the sales it is measur so tblo going to aggregate first and
then combine it so the sum of that going to be 80 so let's do that and the next one we have the date so here it is a
dimension cannot be like aggregated and since we have two different values tblo going to write at the output star and
since tblo going to provide at the output only one value and we have here two values tblo will not decide which
one of them going to be so tblo can to adds a star so what going to happen and the output going to be star I know this
is really not nice but this is how data blending works so as you can see Tableau always try to aggregate the data before
combine it now let's move to the next customer we have John and in the orders we have only one records that's means
nothing going to be aggregated the output going to be exactly the same and then for the customer charge there is no
information over here we will get as well nulls and this is the output of data blending and this is exactly what I
mean with the main differences between joints and blending is when we do the aggregations so in the left joint as you
can see first we combine the row data togethers and afterward we can do aggregations in the visualizations but
in data blending first the data should be aggregated especially from the secondary data source and afterwards the
data going to be combined in tableau [Music] all right so now what are the main
differences between joints and relationships if you are using joints things going to get really static and we
might lose as well a lot of data but if we are using relationships in our data model then we will get more flexibility
and we will not lose any data and now in order to understand this let's check this example where I have prepared two
data sources one with joints and the other with relationships the first one with the orders if I go to the physical
layer you can see we have a lift joint between orders and customers and let's check the second one we have the
relationships we have as well the same tables we have orders and customers and between them there is a relationship and
now if we check our data we can find that there is five customers and in the orders there is only four customers that
did order so if you check over here the customer ID you will not find the ID number five so that means this customer
didn't order anything this is no problem for the relationships but if you go to the joins over here and you check the
data you will see that we don't have a customer ID number five at all in our data so you can check okay we have 1 2 3
four and so on so the customer ID number five is completely disappeared and that's because we have a lift joint
between the orders and the customers so only the matching RADS from the right side is going to be presented at the
final table so that means we lost this customer and if we are at the visualizations let's go over here and
let's say we want to count how many customers do we have in our database so let's drag and drop the customer ID and
let's turn it to measure of count distinct so our data says okay we have four customers if we go to the
relationships let's open another one and switch to the relationships and let's take the customer ID again over here
switch it to a measure and count distinct you will see we didn't lose the data we have five customers in our
database and the relationship is going to give us more correct answers and now you might say okay we can fix this if we
change the type of join so that's right if I go to the data source and then I go to the
joins go to the orders and I just switch this to the right so that mean we going to get all the data from customers and
only matching from the orders let's close this and go back to our sheet number one you will see let me close
this we'll see that we have five customers so with that we have correct answer as well as with the join and here
we come to the next point that things are really not flexible so that means if I'm building visualizations where
sometimes I'm asking how many customers do we have or how how many orders do we have I cannot each time go to the data
source and change the type of joint because once I decide it's lift joint it's going to stay for all the
worksheets as a lift join unless I'm doing full outer join between the two tables and if you are working with big
tables then you will get a very big mer table which going to slows everything down and this is exactly what I mean if
you are using joins you will lose data if you are using left joint or right joint and as well things are really
static with the relationships if we go to the sheet number two here things are more fle ibles because we didn't merge
anything the data stay separated from each others we just describe the relationships between them so if in
worksheets I'm doing analyzes about the customers it will not affect the next visualizations if I'm doing analyzis
about the orders because we didn't lose any data and I don't have to worry do we have left joint or right joint should we
change it and so on so it's more flexible and we will get always correct answers so that's why joints are static
and you might lose data but relationships are more flexible and you will not lose any
data all right guys so there's another issue with the joints if you compare to the relationships sometimes in joints we
might get wrong answers if we are doing calculations on the measures so let's take this example on the customer tables
we have the score so for each customers we have a score and we have those five customers the average of this score is
going to be 625 and now let's sck in tblo the results from joints and relationships all right so now we are at
the relationships and let's take the score and and drop it over here on the text and then let's find the average so
we're going to go over here measures and the average so in relationships we got the correct answer we have 625 and now
let's check the joins we are at the data source of joins I'm going to take the score drag and drop it on the text and
now we're going to switch as well to average and here we got the wrong result 585 so what happened here well the
answer for that is sometime if we merge two tables together we might get duplicates so let's check the data if
you go to the data source I again in the joins if you go to the score we will have duplicates because
some customers have more than one order and that going to result in a lot of duplicates if we merge the customers and
orders and if you do the average you will get the wrong answer as we saw in the results and if we switch to the
relationships and we go to the customers we see at the score over here on the right side there is no duplicates and we
will get the correct answer and that's going to guarantee for us that using relationships we will get correct
answers if you are doing calculations and that's way better than having duplicate in our data we might never get
correct answers from joints and that's why Tableau introduced in 2022 relationships just to fix all those
problems with the joints and they made it as the default method on how to connect
tables all right guys so now we're going to go and compare the four methods on how to combine data in Tableau unions
joints relationships and data blending side by side so let's go the first point is in which page in which layer we can
use the method now both Union and Joints we can create them at the data source page in the physical layer and as well
the relationship we can use it at the data source page but in The Logical layer and finally the data blending
could be used at the visualization level in the worksheet page and the next Point can we use the method in order to
connect tables from different data sources well for Union joints and relationships we cannot do that it
should be done in the same data source but only the data blending could be used in in order to connect tables from
different data sources the next point is after using the methods are the tables going to be merged in unions and Joints
they going to merage the tables and they going to create completely new tables but if you are using relationships and
data blending they will not create anything the next point is about the flexibility if you are going to use
unions and Joints the decisions that you are making at the data source going to affect all the worksheets and the
visualizations but if you are using relationships and data blending you have way more flexibility for example in the
data blending you can decide in each worksheet page now if you are talking about the joint types in joints we have
inner left right and full in the relationships we can have as well exactly the same behavior as joints but
in data blending it is fixed we have only left join and the next point if you ask me to rank these methods I would say
and tblo as well going to say always use relationships and after that comes the data blending it is a really great way
on how to combine tables from different data sources and the flexibility that we have and then the third one I'm going to
say the joins I would not trank Union because it's completely different than the methods of joining relationships and
data blending so always try to go with the relationships and now let's see the big
picture on how those four methods works and let's start with joints they're going to connect two tables at the
physical layer and they're going to create completely new logical table in The Logical layer where it's going to
combine the fields of both tables and then at the visualization layer the data sets going to create query at the data
source and and data source going to get the data from The Logical table and same thing for the union you can create it at
the physical layer of two tables and then going to create as well completely new table where the rows of both tables
going to be combined and at the visualizations table going to send a query to the data source and the data
source going to get the data from The Logical layer and now to the third method of the relationships we have two
tables at The Logical layer and tblo will not combine or create anything we are just describing the relationship
between a and b and at the visualization level tblo going to ask the data source and the data source going to get the
data from the Separate Tables and finally the data blending we have two data sources the first one going to be
called the primary data source the second one is the secondary data source so first T going to send query to the
primary data source and then another query to the secondary data source here it's important that the aggregation
going to happen before the data is combined and we are combining the data at the visualization level using data
blending so as you can see joints and Union happen in the physical layer in The Logical layer we can do
relationships and at the visualization level we can do data blending all right guys so now we're
going to create together two data sources because we have two data sets the big one and the small one and during
that I want to show you how I usually make decisions on when to use which methods so let's
go okay guys so now let's close everything and start from the scratch in order to get the data source correctly
created so let's let's start Tableau public we're going to create now the small data source on top of our small
data set so let's go to the connectors on the left side and click on text file and then it doesn't matter which one
you're going to use let's take the orders open I will delete it anyway in order to explain how I start so
previously I showed you the data model of our data sets we have a star schema where we have facts and dimensions I
always start with the fact table doesn't matter whether you are using star schema or snowflake always start with the fact
table so our fact table is orders so let's just drag and drop it here on the logical layer and then I continue with
the dimensions so we have customers and products so let's start with the customers just drag and drop somewhere
over here and table going to create a relationship between the orders and customers and since we are talking about
two different entities so we have orders and customers I always use relationships between them and now let's check the
relationships whether everything is correct so we go over here on the metadata we see the customer ID from
left the customer ID from right which is correct and now let's go to the performance options I will change only
the cardinality if the quality of our data is bad and we haven't done any data profiling then the best is to leave it
as default so many to many some record matches on the left and on the right but in the data sets we already check that
so we have clean star schema and always on the fact side on the left side over here it going to stay as many and all
the dimensions on the right side like customers it's going to be one because we have usually for example unique
customers or unique products so I will go and and chain that on the right side as one because it is dimension side and
on the fact side it's going to stay as many I will not touch those Integrity stuff so we're going to leave it as it
is and that's it we have now the customers and the orders connected to each
other and now before we continue building our data model we have to check something very important are we working
on the correct data sets in the correct format so now if you go to the orders over here and here we have some few
Fields like the sales quantity discount all those informations should be in number and you can check that by
checking the icons the data type icons and if they are like this hash value over here and green if you click on it t
going to say it is number decimal so if you see it like this number decimal or number then everything is fine but if
you see it as a string for example if you go over here and switch it to string so if you see this field as a string
there is something wrong so if your data is like ABC then you are working with the wrong data set it's not correct so
you should see it like a number so now the question is why it's wrong why it's not correct why Tableau didn't find it
as a number well there is different representations of the decimal separator in decimal numbers some countries like
in Europe we have a comma but in many other countries like in USA in Asia we have a DOT between the decimal number
and the whole number so now for example I'm now in Germany and my data is separated with a dots what going to
happen tblo will not understand this is a decimal number and it going to show it as a string and that's why in the
download link I have prepared two data sets depend on your location the Europe training data sets and the non- training
data sets the Europe training data sets all decimal numbers are separated with comma and for all other countries they
are separated with a DOT for the first downloader so now the question is how to fix it well go and download the correct
training data set there is another way in order to fix it for example now I have the Nur data sets and as you can
see the Discount sales profit everything is wrong everything ABC and strange now some of you thinks okay it's really easy
fix I can go to the data type over here and switch it from string to a number decimal so once I do that what's going
to happen everything going to be null so it will not work because Tableau don't know how to convert those numbers
correctly so let's move it back to a string in order to see the data again there is a fix for that if you go to the
orders over here and then right click on it and let's go to the text file properties so here we have different
properties about the files like the separator here we have it semicolon so TBL detect it correctly but what's more
important than this is the format of the decimal number the local so here we have to choose a local which is matching to
the current format so the current format is a DOT here in this example so what we're going to do we're going to go over
here and search for for example United States and as you can see table going to understand the correct format and
everything going to be changed to a number so the solution either you're going to use the correct data sets or
you can go and configure the properties of each file so I would say you can go and try Unite uned Stat or Germany until
you have the data type number so make sure that in the orders all those informations is the data type number all
right so now let's go and keep building our data model in the data source let's go to the next Dimension we have the
products so all what we're going to do is just drag and drop and then release it TBL going to create another
relationship between them let's check that again so click on that go to the metadata scroll up so Tableau did
automatically find the key for the relationship it is the product ID which is correct and now the same thing we're
going to go to the performance options on the left side on the fact side it's going to stay as many and on the right
side it's going to be one so on the right side we have the dimension it's going to be one you can check that
easily if you click on the product and here check the data you can see the product ID is a unique field there is no
duplicate inside it and we can go and use one if you are not sure just leave it as many to many relationship so let's
go again to the relationship we have it many to one and I'm going to leave it here as some recards matches no problem
and now let's go to the other tables we have here the customers details and here we have two options either we're going
to use relationships or joints so you can go over here and just drag and drop put it near the customers as a
relationship but to be honest in data moding if I have two objects about the same entity so here we have customers
and here another informations about the customers I tend to merge those two tables in one this is different than
talking about the orders and customers they are completely different entities and usually in data warehouses I I
prepare this step in the database or we can stay on Tableau and merge those two tables into one and we can do that using
joints so what I'm going to do I'm just going to remove the customers details away and then we're going to go to the
physical layer inside the customers and then we're going to take the customers details and drop it over here and tblo
as default going to leave it as inner joint but to be honest the customer table is for me the main table about the
customers and customer details is like secondary table so in order to not lose anything from the left side I'm going to
change the type of joint to left join so let's do that I'm going to click on the icon and then select left join then we
can check the results well the main thing that we don't get dcat or we don't lose any customers so as you can see the
output we have our five customers there is no duplicates and we didn't lose anything so let's go back to the logical
layer and just going to close this so as you can see we have list tables and we have one entity called customers we
don't have a lot of tables and I usually do that if we have a lot of tables about the same topic and now let's go to the
next table we have the order archieve and here we have the same situation we have two tables describing the same
entity the orders but of course we can connect it as relationships to the orders but again I like to minimize the
number of tables that I'm dealing with and I'm going to go and merge those two tables together and so here we have
again two options unions or joins if the tables has exactly the same number of columns and the same data types then we
can use Union in order to do that we have to do data profiling so either you open the CSV file and compare them
together or we can go over here there's like small icon like a table and if you click on it tblo going to show you a
symbol of data in order to do data profiling and to understand the content of this table so let's just make it
bigger so we have the order date shipping date customer ID product ID and as well the unit price and so on and we
can compare it to the orders over here and let's just make it bigger and we can find exactly the same number of
fields the same contents the same data types so that means we can go and do Union between them so in order to do
that I'm just going to close this and go to the physical layer inside the orders I like to drag and drop just beneath it
over here and now you can see we have a union let's check that on the right side in
the table name so we have orders and we have orders archieve so with that we combined both of the tables in one
logical table so let's close this and as you can see we have the icon that there is inside it a Union and with that we
have only three tables instead of having five tables it is just easier at the visualizations to deal with three tables
instead of five tables and the data model is much easier to understand and to explain and with that we have
connected all the CSV files together but we still have one file the Json file product prices sadly we cannot connect
it with the others in the same data source because it is different file type but we still can connect it to them if
we create a second data source and use data blending and now that sets we have our fact table and the Dimension we're
going to give it a name so I'm going to call it small data source and now you can pause the video
and go and create the big data source and if we are done I'm going to go and create the big data source so I'm
going to go over here new data source going to click on the text file then I will just go back to the big one here we
have only the three so we start with the orders always we start with the fact table and then we take the dimensions
let's take the customers customers I already checked all those IDs they are unique so I can
go to the relationships over here and change it to one on the right side and on the fact side it's going to stay as
many the same we going to do for the products drag and drop and all the ideas of the products
are unique so we can go to the performance option just to make sure we selected the relationship and select one
so that's it I'm just going to call it big data source so now in order not to lose those data sources in Tableau
public we have to publish to our public account so I will go and do that we're going to go to the sheet over here and
let's just take something like the customers drag and drop on the rows and that's it I will just go over here and
publish it save to tblo public and I have to sign in I'm going to call it data
sources then save and now it start publishing to our profile so that's it if you want to
download the file you can go over here and download Tableau workbook all right guys so with this we have created two
data sources on top of our data sets and we can to use them in the whole tutorial all right guys so with that you have
learned everything about the Tableau Data modeling in data sources and how to combine tables using the four methods
and in the next section we will start talking about the metadata in Tableau we will learn there are many important
tableau concepts for data visualizations the metadata of Tableau understanding the Tableau metadata
Concepts like data types measures Dimensions discrete continuous is very important in order to build a correct
data visualizations in Tableau and as well going to help you to understand how Tableau works with your data so first
I'm going to introduce you to the metadata in Tableau to learn what happens to your data once you connect it
to Tableau next we're going to dive into all data types in Tableau like integer string date and so on and after that
we're going to learn about the data type rules like the geographic Rule and the image rule and after that we're going to
cover very important Concept in Tableau we have Dimensions measures discrete and continuous and of course in order to
understand the differences between them we're going to compare them side by side in order to understand the big picture
so now let's start with the first topic where we can have an overview of the basic concepts of metadata in Tableau so
now let's go [Music] all right so now we're going to have a
quick introduction to the Tableau metadata in the data sources in order to understand what's going to happen to our
data once we connect it to Tableau after connecting our data to Tableau and building the data model in the data
sources the next step is to check the metadata of the tables and the fields because once you connect your data to
Tableau Tableau can to start analyzing the content of your data to make assumptions about the types and roles of
each field in the data source so tblo going to assign each field to data types like integer string date and so on data
types gives us informations about the kind of data stored inside our data sets this piece of information is very
helpful for Tableau in order to understand how to deal with your data which rules operations calculations can
be performed and one more thing that Tableau going to do it's going to assign each field to a role these roles going
to help Tableau building the visualizations so the first set of roles we have dimensions and measures
Dimension fields define the level of details of the view and the fields with the RO measure going to be used for
aggregations in The View and we have another set of roles we have discrete and continuous these rules can help
Tableau by plotting the visuals so discrete Fields can break the view to separate values and the fields with the
continuous rules going to plot unbroken chain and connected values in the view and I call all those informations about
your field as a metad DAT in the Tableau Data Source one more thing that I want to tell you is that those assumptions
that Tableau makes about your field is correct around 90% so that means there is a possibility that those assumptions
from Tableau are wrong that's why it's very important after you build the data model is to have a double check on the
metadata to check that all the informations are assigned correctly otherwise you're going to have bad
quality and bad results at the visualizations all right so next we're going to do a deep dive into these
important Concepts in order to understand them and the differences between them
[Music] all right so we can find data types not only in Tableau but in all programming
languages but they don't support exactly the same data types and that's why if you are learning new programming
language or an application like Tableau it's very important to understand which data types they support and now the
question is what is a data type the data type give us information about the kind of information stored inside our data
and this piece of information is very important for programming languages and applications like Tapo in order to
understand how to deal with your data which rules operations and calculations could be performed on top of your data
now if you look closely to our data you can see that each field in our data source must be assigned to a small icon
or a symol those icons indicates the data types of each field and now one more thing once we
connect our data to Tableau Tableau going to analyze our data in order to assign automatically the correct data
type to our Fields well most of the times Tableau does it correctly but sometimes things goes wrong or you want
to change the data type of specific field this is really easy either you can do it on the worksheet page or at the
data source page you will get exactly the same effects so let's go to the data source page and let's go to the orders
and click on the icon over here you can see it's number Hall we can change it to a string so what we're going to do we
just going to click on the string and that's it we just changed the data type of the order ID but let's say we want to
change it back as Tableau did it at the start what we're going to do we're going to go to the icon over here again and
then we go to the default and it's back to the original data type that Tableau did assign at the start and here one
more thing to notice is that the data types are really sensitive in the joints and the relationships for example if we
go to this relationship over here between the orders and the customers the key is the customer ID those keys should
have exactly the same data type so let's say we go to the orders and let's change the customer ID from number to string so
we're going to go to the string over here and we change it and immediately you can see at the data model the
relationship between the orders and customers is now broken and you can see at the tool tip it going of says type
mismatch between the customer ID the string and the customer ID number so as you can see now Tableau is very
sensitive with the data type of the key whether you are using relationships joints data blending doesn't matter they
should have exactly the same data type so now in order to correct it as you can see we don't have anym the data review
the data grid so how we can change now the data type we're going to go to the metadata grid and we're going to do the
same thing we're going to go to the customer ID just click on the data type icon and change it back to defaults or
to number so I'm just going to click on default and tblo going to be happy now and the tables are related again and the
third way to change the data types you can go to the worksheet page and same thing over here you can go to the icons
and change the data type so as you can see it's really easy in Tableau we have bunch of
different data types that's we're going to cover in this tutorial and I group them into three categories first we have
basic main six data types we have the number hole number decimal string date date and time and bullion and the second
group we have roles we have Geographic roles and image roles and the last group we have Advanced Data types like group
cluster group bins and set and this group contains special data types that's introduced from Tableau for data
visualizations and they are specially made in order to organize our data in this tutorial we're going to focus on
the first two groups the basic and the role and for the Advanced Data types I'm going to dedicate another full tutorial
just speaking about them all right so now let's start with the first group the basic data types where we're going to do
deep dives into each type in order to understand them so let's go all right so now we're going to talk
about the data type number if our data contains only number nothing else it contains digits from 0 to 9 then we can
call it a number data type and it's very important to understand that numbers cannot contain any characters for
example let's say that we have the following phone number in our data this type of data we cannot call it a number
because it contains characters like we have the minus we have the plus because the number data type can only have
digits from 0 to 9 and now if we remove those characters from the phone number then it's going to look like this and
only now we can give it the data type number and in Tableau the data type number has this icon it's like hash and
for numbers we have two data types in Tableau we have number ho and number decimal so what is the difference
between them you know in math a positive or negative number could be splitted by dots the first part we call it a whole
number and the second part we call it a decimal so if your numbers does not include decimal dots or any fractions
then we can call it a whole number like 3 minus 100 zero and so on but if your number contain dots and fractions then
we call it a decimal number like 2.4 or 3.99 and here you need to be careful which one you are using especially if
you are making calculations in Tableau for example if you want to divide two numbers like 1 divide by two if the
output field has the data type whole number then the result going to be zero but if it has the data type number
decimal then the result can be correct 0.5 and this is exactly the difference between those two data types all right
so now let's check our fields in Tableau to find out which one has the data type number and I would say let's check the
orders over here and you can see we have the order ID customer ID product ID by just checking them you can find that all
of them are numbers they don't have characters and they don't have fractions so that means they should have the data
type number hole as you can see all of them is number hole let's check another fields on the right side we have here
sales we have discount profit and as you can see they have fractions so those numbers should be a number decimal so
let's check that you can see TBL did automatically fig fig out that those numbers are number decimal but for the
quantity it's whole because we don't have here any fractions so that says everything is
fine all right now we're going to talk about the data type string the string data type is one of the most widely used
data type in all programming languages a string data type is a sequence of characters and it could include anything
like letters numbers bases and any other type of characters and you can think of a string as a plain text and any field
in our data source could be a string so string is like a default data type and it has no rules or whatever like the
other data types so that means you can convert any fields in your data source to a string data type without any
problem and Tableau as well uses the string data type when it couldn't find any suitable other data type for your
Fields so now let's check in our data sets where we can find Fields with the data type string so let's check first
the products over here you can see we have here two strings the product name and the category in the product name we
have characters we have spaces we have numbers so those are the data type string let's check the customers over
here we have the first name last name both of them are string but now you might notice or ask you know what we
have City and Country both of them contains like characters why don't we have the icon of ABC is it like string
well the answer is yes because if you just click on the icon you can see that tblo did assign it to a string but here
the difference is that they have an extra role we have the geographical role and you can see tablet did assign it to
a country and here tblo going to give it another icon just to indicate that this field has a geographic role but the
basic the main data type for that is a string and the same is for the city okay now we're going to talk about
one of the most confusing data type it is the date if your field stores informations about the calendar data
then this field is going to has the data type dates and dates have very different formats in different countries for
example in Germany we have the following date format you see we use dots instead of slashes but date in the international
formats follow another rule where the date can to split it by a minus and in the world there are many many different
formats so those dates follow specific formats and we describe it with the following code for example for the
international formats we have this code it's going to start with the year and the year has four digits that's why we
have four times y then we have a minus and two digits for the moners so we have mm minus two digits for the day d d so
there is like a code for each part of the dates we have the day months year weeks and so on in this table I'm going
to leave the link on the description you can find all those codes and the descriptions of That So with that you
can customize the date format as it suits you and don't worry about it Tableau understands almost all date
formats that we have and in our data we could have not only the calendar data but also informations about the time
then we have in Tableau another data type for that we call it date and time and in programming languages or
databases you might heard already about the timestamp but in Tableau we call it date and time so it might look like this
we have the date then space and then afterwards we have informations about the hour the minute and the seconds and
like the dates it could has as well different formats you could have the millisecs or the time zone and many
other stuff so here we have again a table of all the codes for the time informations you can find it as well in
the same link all right so now let's check our data to find out which Fields has the data type date usually in Star
schema data model all the dates are placed at the fact table and our fact table is the orders so let's check that
you can see we have two Fields with the data type icon dates we have the shipping date and the order date and
it's not date and time because we don't have in the data informations about the time so both fields are dates we can
check here and as well here and in the other tables products and customers they don't have any dates or times because
they are dimensions they are not events and usually don't don't have any informations about the date all right so
now let's go back to our orders to our two fields and as you can see the format here is that they are splitted with
slashes let's say that you don't want this format you want something else so now how we can change the date format in
Tableau in order to do that we have to go to the worksheet page so let's go to the worksheet page over here and now you
have to decide something do I want to change the date formats for the whole workbook for the old visualizations so
that means you are changing the default format of the date or you want to change the format only for this view only for
one visualization so let me show you how you can do both so now let's put something at our view I'm going to take
the order ID drag and drop it over here and let's work with the order date so I'm going to drag and drop this on the
text table going to show it as a year so I want the exact date in order to see the format so as you can see our date
has the following format and now I want to change the default date format for the whole workbook so in order to do
that we're going to go to the left side to the order date right click then we go to the default properties and here you
can find the date formats so if you click on that automatic it is what Tableau did figure out at the start and
then we have some predefined formats from Tableau what is interesting is at the end we have custom so our new format
for the date going to split with the dots and the year going to have only two digits so the code format going to be
like this DD for day then dots mm for month and for the year we're going to have only two digits that's going to be
y y twice let's hit okay and as you can see tblo did change the date format in tblo blow so now let's go and duplicate
this worksheet over here by right clicking on it and then duplicate as you can see in the next worksheet as well we
have exactly the same format that we defined so this means that the format that we defined is a default now for the
whole workbook but now let's say that I want to change it only locally at one visualization and I don't want to change
the default format for the date so let's duplicate that as well once again and now instead of going to the
left side we're going to stay at the view and we're going to go to our field right click on it and then we go to this
one here format once you do this on the left side the datab ban going to switch to the format Pane and over here on the
left side you can see dates so if you click on that we're going to get exactly the same stuff over here those are the
predefined from Tau we have the automatic at the top and at the bottom we have the custom so now let's choose
one of those predefined I'm going to take the week and the year so let's click on that so as you can see tblo did
change the date format in this View and now interesting to check the other sheets whether the date format did
change so let's go back to the previous sheets and as you can see they stayed at the default format of the date so with
this you learned how to customize the format of the date for specific view or for the whole workbook but now I want to
change the date format as before so in order to do that so I'm going to go over here close this formats then go to the
order date again right click default properties date format and then we just click on the automatic and hit okay so
as you can see we have again the same old date for so that's it this is how you can work
with a data type date all right now we're going to talk about the last data type in the basic
category the buan data type the buan data type represent a fields that has only two values true or false it's like
the language of computer we have only one and zero and this data type is often used in the output of a condition or
logic so for example if I ask you do you like this video so far so the answer is going to be yes or no if you like this
video please give it a like so the answer for this question can to has the data type bullion either yes or no true
or false and no any other values and don't forget to subscribe so the bullan data types has many use cases for
example control the workflow of something if the output is true then do something if false then do something
else all right so now let's check whether we can find any bullan data type in our orders we can check over here we
don't have any Pion data type and the customers as well nothing and in the products well we don't have
any fi with the Boolean data type well usually data type Boolean going to be add once we use conditions in taow and
once we create new calculated fields and now to create the calculated field we're going to go to the worksheet page so
we're going to go sheet number one and now make sure to select the small data source then we go to this small icon
over here and now we select create calculated field so let's click on that we will get a new window to write our
expression or our condition I'm going to give it the name of logic 400 and now what we going to check or what is our
condition if the sales is smaller than 400 then should be true otherwise going to be false the logic is very simple so
here we're going to find the sales smaller than 400 and that's it if the sales is smaller than 400 it's going
to be true otherwise going to be false let's click okay and once you do that you can find on the left side we have a
new field called logic 400 and it has the data type volum the output has only two values true and false so let's
validate that I'm just going to drag and drop this on The View over here and as you can see we have only false and true
and let's check whether the logic is working so we're going to take the order ID and just put it before it and now we
need the sales so we're going to take the sales drag and drop it here on the APC and here you can see for example the
first order it is smaller than 400 that means the logic is true which is correct and then the next one it is above 400
it's false and so on so we can see if the field has only two values true and false then the data type going to be
Boolean and we usually use it as an output of a condition and the Boolean data type has a lot of use cases for
example if you want to filter our data anything above 400 we don't want to see it in our visualizations so what we can
do we can use the logic in the filter just drag and drop that on the filters and we're going to select only the true
so I'm going to unmark the false and then hit okay and as you can see the result going to show only the orders
with the sales less than 400 and with that we just filtered our data very easily
all right so with that we have covered the basic six data types in Tableau so now let's do a quick recap we have the
number hole is for fi that stores only numbers without characters and those numbers are without fractions or decimal
dots next the number decimal is as well for fields that have only numbers without characters but those numbers
could have fractions or decimal dots string is a sequence of any characters it could be numbers letters special
characters or spaces and then we have date date is for fields that stores informations about the calendar dates
next we have the date and time is as well for fields that stores informations about the calendar and as well about the
time and it has as well specific format and the last time we have the bullion it can store only two values false or true
and we usually use it for conditions okay guys so the first role that we're going to talk about is the
geographic rule if you have in your data field that contains location informations or geographical areas then
you can assign it to a geographical role in Tableau based on the type of the location such as City Country postal
code and so on assigning this extra role can to help Tableau to plot your data correctly if you are using map
visualizations in Tableau there are over 12 Geographic roles but I think the most important ones are country City and zip
code now let's check our data but first some coffee let's go all right back to our
data source let's go to the customers table there we have some informations about the location of the customers and
here we have three fields we have country City and postal code and now in order to check the geographic role just
click on the icon over here on the data type and again here it's very important to understands each field must have a
basic data type so for example the postal code is a number hole and then we assign an extra role for it so having
the geographic role will not remove the number data type so now let's check the geographic Ro over here and you can see
that Tableau didn't assign it to anything so it stays here none and this is a ZIP code or post code so we're
going to correct that we're going to just click on this over here to assign a geographic role and you can see the icon
did change so with that we have the data type number and we assigned a geographic Ro for it let's check the others so this
should be a city so let's click over here the basic data type is a string because we have characters and let's
check the geographic role tblo did it correctly we have it as a city so that is correct let's go to the country over
here we have it as a string and then the geographic role is country so with that
we have all location informations assigned correctly to the geographic role and we can start building a map
visualizations in Tableau let me show you an example let's go to the sheet number one over here and what we can do
we can go to the customers over here and let's take the location informations so let's take the country the city and
let's have one metric so I'm going to take the sales drag and drop it over here on the APC so as you can see it's
only a table we you want to switch it to a map in order to do that go to the show me over here and then click on the map
so you can see Tableau did correctly plot our data let me just close it and assign for each country The Matrix and
this is done because we assigned our data to a geographic role all right so now let's talk about
the other one we have the image roll this is brand new tblo just introduced that in 2022 so in princip if your field
stores a URLs pointing to IM then you can assign this field to image Ru with the URL to show the images in
the visualizations and Tableau have here some requirements so the first one table supports only those three image
extensions and the URL should begin with the HTTP or https and the third requirement the maximum number of images
in each field is 500 and then we have the image size it should be less than 128 kiloby but though things might
change in the time since it's completely new feature in Tableau and I think the most Ed Case for this is to show the
product images in your visualizations all right so now let's see an example in Tableau about the image role in our data
sets I have prepared some URLs inside the table products but only in the small data sets so let's check that if you go
to the products over here we have a field called Product images and here we have URLs pointing to images in my
website so now let's check the data type over here it is a data type string this is the basic one because a URL is a
sequence of characters and now we can add on top of this basic data type an image image roll and it's really easy we
just go over here to the image roll and we click on the URL so let's do that and with that we have a new icon indicates
that this field has the role of image so let's check the data we're going to go to the sheet number one then we go to
the products make sure we are selecting the small data source then we go to the products image just drag and drop over
here and as you can see now we have some images about the products but two of them are broken and I think it's still
bugging at the disktop version of Tableau public because if we publish now to table public in the web we're going
to have all the icons correctly so now we can go and grab another field let's take the sales drag and drop it over
here and with that we have a nice images to The Matrix let's go and publish that in tblo public I'm going to call it view
with image let's save and as you can see now in tblo public we have all icons nothing is
broken so I think if you are building dashboards about the products it's really nice to show the image of the
product instead of the names it's just more catchy have images inside the visualizations dimensions and measures
in Tableau so once we connect our data to Tableau Tableau going to analyze our data in order to assign each of our
fields to either a dimension or measure this kind of metadata going to help Tableau to blot our visualizations all
right so now the question is what is dimensions and measures well Tableau didn't invent the concept of dimensions
and measures it is an old concept of Pi and now we get going to have a quick origin story if you learn the concepts
of data warehousing and business intelligence you might already know that the core concept is the
multi-dimensional olab online analytical processing so the concept says if you want to answer the business questions or
do data analyzes first we have to build a data model that has the shape of a cube with multi-dimensions it's
something like this Cube and each Cube has two informations first we have the dimensions of the cube and the second
informations we have those cells those cells can store informations like data numbers and we call it measures so each
Cube has two informations the dimensions and the cells the measures and now let's have an example we have the cube of
sales and it has three dimensions the First Dimension is the locations and inside the locations we have three
members USA France and Germany those three values are the member of the dimensional location and we have another
dimension called time and it has three members in the dimension January February and March and the third
Dimension we have the categories and now inside the sales of the cube we have the measure sales so now our cube is ready
with the dimensions and measure and we can start answering the business questions for example find the total
sales in USA so what going to happen we can select the dimensional location and filter the dimension to have only the
member USA this operation in the cube we call it slicing the cube and then we're going to aggregate the measure and we
will get the total sales of 120 and if you have Cube we can do multiple operations like slicing dicing roll up
drill down and beot so if you have such a cube we can do data analysis and find fast answers to the business questions
so now to summarize Dimensions contain qualitative values they usually describe something like the product name the
product category customer location and we use Dimensions to categorize filter and show the level of details and in the
other hand we have the measures they contain numeric quantitative values that can be measured like the name says and
the measures unlike the dimensions they can be aggregated all right so this might be
still confusing and if you say you know what if I look to my data how do I decide whether it's a dimension or a
measure so here is my decision- making process first I check the data type of the field whether it is a number if the
answer is no then this field is a dimension but if the answer is yes then we going to ask the next question does
it make sense to aggregate the values of the field like doing the sum calculation on the Val values or finding the average
value if the answer is yes then it is a measure but if the answer is no then it is a dimension so what this means all
non-numeric fields are dimensions but not all numeric fields are measures that really depends on the questions whether
it makes sense to aggregate the values if yes then it is a measure if no then it's Dimension okay so now let's
practice in order to understand the concept of dimensions and measures and how they work we will check our data
sets and we going to assign each field to either Dimension or measure we going to do the table customers together and
then you can go and pause the video in order to do the products and the orders and then at the end we're going to check
the result together so let's go we're going to start with the first field the customer ID the customer ID is a number
so we cannot say it is automatically a dimension we're going to jump to the next question now does it make sense to
aggregate it well we have here to understand that the customer ID is a unique identifier for the customers for
example Maria has the customer ID number one Martin has four and now if we sum all those values we're going to get the
value of 15 or if we do the average we're going to get the value of three those values don't make any sense
because we use the customer ID only to identify the customers and I don't think that we will be in situation where we
have to find the average of the unique identifiers so since it makes no sense this field is a dimension and with that
we can assign the customer ID to a dimension now let's go to the next one it is much easier because we have here
the first name and it is not numeric so it is automatically Dimension the same goes for the last name it is as well a
string it is not a number all right so now let's move to the next one we have the post code or the ZIP code it is a
number so we can to ask the question does it make sense to do aggregation here well I don't think there will be
situation where we have to find the sum of the postcode or to find the average of it so that means it is here again
it's a number but it is a dimension so let's assign the value for that and then the next one it is easy so we have the
city and the country both of those values are string so it is automatically a dimension so let's assign it
again okay so let's move to the last field we have the score here it's again a number so we going to ask the question
does it make sense here to do aggregations well the answer is yes it really makes sense to find the average
of the score that's why we're going to map it to a measure so on the table customers we have six dimensions and
only one measure and now you can go and pause the video in order to practice with the table orders and as well with
the products all right so now let's check the results as you can see in the table
orders we have a lot of measures because it is a fact table and fact tables in the star schema is the central place for
the measures so this is very normal so let's check the fields we have the order ID customer ID product ID it is like the
customer ID those are identifiers and it doesn't make sense to aggregate it so that's why we have it as Dimensions the
order date and shipping date those informations are not numeric and that means it is dimension and then we have
all those informations the sales quantity discount profit unit prices all those fields are numbers and here it
makes sense to do aggregations like the sum or the average so we're going to use the orders the fact table if we need any
measure let's go to the next one to the products and here this one is easy the product ID is like again the identifier
it doesn't make sense to do any aggregations we can have it as Dimensions product name and category
both of those informations are string they are non- numeric and that's why they are dimensions so I hope with this
you have understood how I usually do it by just looking at the data we could decide whether it's a dimension or
measure all right so now back to Tableau and the first question is where do I find in
Tableau whether my fields are measures or Dimensions well there is no icons for dimensions and measures and as well we
cannot check that at the data source page in order to check the dimensions and measures we have to go to the
worksheet page so let's go to sheet number one and then we're going to go to the datab ban on the left side over here
let's open any table for example the orders and now if you look closely to the table orders you will find like fine
gray horizontal line which splits the fields of the orders into two groups the fields above the line they are the
dimensions and the fields below the line they are the measures so for example we have the customer ID the order dates
order ID product ID and so on those fields are dimensions in Tableau and the fields below the line the discounts the
quantity sales and so on those fields are measures and you can find this splitter this horizontal line in each
table so if you go to the customers over here you will see again the same line that splits Dimensions from measures and
and the same if you go to the products scroll down we have again the same line and one more thing that you might
already noticed let me just close those tables that outside the table there is as well a horizontal line sometimes in
Tableau we create fields that doesn't belong to any tables and Tableau going to put it just outside of the tables
it's like Global fields and for that we need as well a splitter to split the fields to dimensions and measures okay
so now let's go back to the orders and now you might say you know what we don't need this horizontal line to identify
whether the field is dimensional or measure and now if the field has the color of blue then it's Dimension and if
the field has the color of green then it is measure well this is exactly where most of Tableau developers get confused
and things gets mixed up between Dimensions measures and discrete continuous and to be honest I was
thinking the same at the start until I found out that the color of the fields indicates whether the field is discrete
or continuous we're going to talk about this concept in the next tutorial don't worry about that so the color does not
indicate whether the field is dimensional or measure but the position of the fields whether it's above the
line or below the line and let me show you quickly something let's take any Fields over here the product ID let's
just drag it little bit and now table going to Mark the horizontal line with orange and going to show you okay
anything above is dimension and anything below is measures so tblo shows that as well all right so now to the next
question how do I change a fields from Dimension to measure and vice versa and here you have two options either you're
going to do it globally for the whole work for all the views or you might do the change locally in one individual
view so let's see how we can do that let's start with the first one where we're going to do the change for the
whole workbook for All Views so globally we're going to go for example let's take the order ID over here just right click
on it and then we go over here convert to measure so let's click on that and as you can see the field order ID just
jumped from above the line to below the line as a measure and now if you want to change it back to Dimension just right
click on it and then convert to Dimension so that's it it's really easy and now let's see how we can do the
change locally at one view without affecting the whole workbook so let's take again the order ID drag and drop it
over here and here we're going to right click on it on The View and then we're going to go to the measures we're going
to convert it to a measure currently it is a dimension so let's go to the measures and we have to select one of
those calculations so let's take for example the sum and now as you can see the order ID only for this view is a
measure but the order ID on the left side for the whole workbook it stays as Dimension and that's it this is really
easy how you can convert between measures and Dimensions all right so now let's have
an examples in Tableau in order to understand the main purpose of measures and dimensions so let's go to the orders
on the left side over here in the small data source and let's take one measure the sales we just going to drag and drop
it on the text over here and as you can see tblo going to start immediately doing aggregations on the measures so
now if we check the data we have only one number this is the total sales that we have in our data set and now we are
at the top level of details where everything is aggregated in only one number and now we have to add more
informations in order to understand this number and in order to do that we're going to use Dimensions so for example
let's go to the products over here and let's take the category so I'm just going to drag and drop the category over
here and as you can see now the dimension is splitting our measure into two rows so that means we have now one
level lower of details than the top aggregation and now let's take another dimension we're going to take the
product name so let's just drag and drop it over here near the category and as you can see using this Dimension going
to give us different level of details about the sales than the First Dimension the category so what happened we just
moved with the details One More Level beneath that and now let's take Third Dimension we're going to take now the
order ID from the orders so just drag and drop it near the product name and now as you can see this Dimension going
to bring us to the lowest level of details where the aggregation of the measure is exactly the same original
value and as you can see the dimensions Define the level of details in our views and each Dimension can to take us to
different levels of details and always if you want to go to the top level of details you have to remove all
dimensions and only have the measure so as you can see as we are removing those Dimensions we are going to the top level
of details another nice way to show that if we go to the tree map visualization so let me just go back over here to have
one dimension let's go to show me and then click on the tree so now you can see our data is splited to only two
details so now as we add Dimensions let's take again the product name over here drag and drop it on the label you
can see the view split it to more details and if we go to the lowest level if you take the order ID again over here
to the label we can see the view is splitted furthermore and now I'm going to tell
you a small secret if you follow it you can generate hundreds of reports even if you have small data sets if you combine
any measure with any Dimension you will be creating a new view or new reports with a title following this pattern
measure by dimension for example sales by product profit by category Quant entity by country so if you follow this
pattern you can generate endless amounts of reports and Views in Tableau all right so now if you count the dimensions
and measures in our small data sets we have around 16 dimensions and 10 measures so that means if you follow
this rule you can generate around 160 views and reports so even we have small data sets we can
generate huge amounts of views and reports so as you can see in the visualizations if we combine both of
them we're going to have sales by order date sales by shipping date sales by country and so on all right so now let
me just show you how we build usually reports in Tableau using dimensions and measures we're going to work now with
only one measure the sales and we're going to make dashboards about it so let's stay at the small data source and
we're going to take the sales from the orders let's just drag and drop it somewhere at the rows and now the
dimension going to be the product name so let's take the product name from the products let's drag and drop it over
here so that's it now we have to call it sales by product so let's just rename the sheet over here right click on it
and rename sales by product all right so now we're going to create another one using the
same measure but different dimension so what we're going to do we're going to just going to go and duplicate it right
click on it and duplicate we're going to have now the sales by category I'm just going to rename it again and let's call
it sales by category and now we're going to remove the product name from here so just drag and drop it somewhere at the
white space and then we go again to the products drag and drop the category on the columns and now we're going to use
different visualiz ations so I'm going to go to the show me over here and let's use the pie chart so click on that all
right so now we have like a pie chart but I would like to show the values so go to the label over here click on it
and click on this Mark show Mark labels in order to show some values so that's it this is our second one all right so
now we're going to create the third one with another dimension we're going to take the order dates but we're going to
show only the months so we're going to go over here and duplicate it again let's just rename it so I'm going to
call it sales by month so we will go now and remove the category just drop it here and then let's take the order date
drag and drop it on the columns we're going to switch the visualizations to power so I'm going to click on this over
here on the bars so as you can see here table going to show the years of the order date we want to have it as a month
so we have to switch dots just right click on the dimension and then over here just select the month so let's do
that let me just close the show me over here and then let's add some labels all right so that's what it for
this view let's make the last one we're going to make sales by country so let's duplicate this again and we're going to
call it sales by country and then we're going to remove the dimension order dates and then we're
going to take the dimension country so just drag and drop it on the rows so now since we have the country we can change
it to a map so let's do that we go to the show me over here and then select the map click on that all right so now
we have a map showing the sales by country all right so now we have those four reports or sheets we can build now
a dashboard in order to create a new dashboard so we're going to go to this icon over here click on it and before we
start I'm just going to give it a name so let's call it sales
dashboard all right okay and now we're going to go and drag and drop all the sheets so we're going to start first
with the country so let's just drop it here in the middle and then we're going to take the category just beneath it
then the product beside it let's resize a little bit to the left and then we're going to
take the last one the Mones and put it over here and as you can see with just four dimensions and one measure we were
able to make a dashboard about the sales and just following this small rule sales by country sales by category sales by
product and sales by month so always measure by Dimension and now it's really easy to train just go and pick another
measure with different dimensions and build different dashboards all right so now let's have a
quick summary where we're going to compare both dimensions and measures side by side in order to understand the
differences between them let's start with the definition dimensions are fills that contains descript values and
measures are fields that contains quantitive numeric values for example we have Dimensions like product category
country and customer ID and in the other hand we have measures like sales profit and quantity the next point is about
aggregating Dimensions cannot be aggregated as each member of the dimension is unique measures however can
be aggregated using functions like sum average minan Max and so on for example you can calculate the total sales for
specific product category moving on to the data types all different data types can be used as Dimensions like string
date Boolean and even numbers like we have learned the customer ID but only the fields with the data type number can
be used as a measure the next point is about the role of analyzes dimensions are typically used for grouping
filtering and organizing your data and measures in the other hands are used for calculations and numeric analyzes and
the final point is about the granularity dimensions Define the level of details of the data and the granularity of
measures on the other hands determines the quantity being measured so this are the main differences between dimensions
and measures all right guys so now we're going to talk about discrete and
continuous here again once we connect our data to Tableau Tableau going to analyze our data in order to make
assumptions where it's going to map each field to either discrete or continuous discrete and continuous are metadata
informations that's going to impact on what type of visualizations that you can create as well as how they will look
like so now in order to understand the concept behind them we're going to compare both discrete and continuous and
first we're going to start with the definition so this concept comes from math and they say discret values are
always separated disconnected distinct values and continuous values are exactly the opposite it's like connected value a
serious or unbroken chain of data without any interruptions so let's have an an example think of discret as you
are counting from 0 to 10 so you start with 0 1 2 3 and so on so that means between 0 and 10 we have exactly 11
distinct values but with the continuous values we have like real numbers which means between 0 and 10 we have infinite
number of real numbers so for example we have 1.2 1.3 1.4 and so on so with discrets we have distinct values and
with continuous we have a range of infinite values between start and end once I read about the discrete and
continuous and the following analogy stick in my head think about the discrete values as IL legal pieces so
you can take them apart and you can work with each piece differently and independently so you can move them
around and analyze them in different orders and now think of continuous as a roll of yarn and now when you unroll the
yarn you will not get different pieces you will just see more of the yarn so you will just get a longer piece of the
same string all right so discrete values are separated distinct values and continuous values are unbroken chain of
data without any interruptions all right so now let's move to the next point we have the colors in Tableau the discrete
fields are the blue pills and the continuous fields are the green pills so let's see in table what this
means all right so now as usual the first question is how do I know whether my fields are discrete or continuous
well it's like the dimensions and measures we cannot check that at the data source page we have to switch to
the worksheet page so let's do that we're going to go over here and now now it's really easy so now as you hover
your mouth on those fields you will see we have only two colors the blue and the green and you can see those colors as
well on the data type icons so we have icons with green and icons with blue the fields with the blue color like for
example the customer ID first name order date and so on those fields are discrete fields and the fields with the green
color like Discount sales unit price score and so on those fields are the continuous fields and here exactly comes
the confusion where a lot of Tableau developer think that the blue indicates for dimensions and the green indicates
for measures well that's wrong those colors to indicate whether it's discrete and continuous so now you know
that so let's start with the first one where we're going to change the role of field globally for the whole workbook so
in order to do that we're going to go to the data Bane on the left side and as you can see here for example the sales
in the orders is green pill that means it's continuous field and as well it is a measure so let's say that we want now
to switch it to a discrete field so in order to do that right click on the field and here we have convert to
discrete it's really easy so let's click on that and now if you check again the sales we have it now as a blue pill so
that means now it is a discrete field so if you check the others all of them are continuous measures but only the sales
is a discrete measure and this change is done globally so if you go to another sheet the sales going to still as a
discrete field so now if you want to switch between discrete to continuous all what you're going to do is right
click on it and here we have again the same option we're going to convert it to continuous so once we click that it's
going to go back to the green pill so that's it it's really easy now we're going to learn how to switch between
discrete and continuous locally for only one view all right so let's build a view we're going to drag and drop the sales
on the columns and let's take a dimension for example the category drag and drop it on the rows and now we want
to switch the sales from continuous to discrete only for this view so what we're going to do we're going to go to
the sales over here right click on it and as you can see the current roll is continuous as Tableau market for us here
or you can see it from the green pill all what you have to do is to select discret so let's go and do that and now
the field sales is discrete for this view as you can see it's blue pill but if you go to the data Bane on the left
side the sale stays as continuous with the color of green so that's how you can do it locally for only one view so for
example if you go back to another worksheet and take the sales the sales going to be a continuous measure so
that's it this is how you can switch between discrete and continuous Fields locally for only one
view all right so now let's move to the next point we have filters in Tableau the discrete field going to create a
filter with distinct values but the continuous field going to create a filter with range values all right so
now let's have an example in order to understand what I mean with those filters and now we're going to work with
the big data source because we need more data in order to understand this all right so now let's switch to the big
data source just click on it and then let's take the sales drag and drop it over here and then we're going to take
from the products the subcategory so drag and drop it on the rows so now we have the sales by the subcategory and
now if we want to go and filter those values we can go and put the subcategory in the filters and don't forget that the
subcategory is a discrete field so let's just drag and drop it on the filters and see what going to happen and now in the
new window as you can see over here Tableau listed all distinct values inside the subcategory and now here with
those discrete values we can make decisions individually so we can include some stuff or remove others so let's
just do do I'm just doing this randomly and click okay and that's it so this is how the filter in tblo going to react if
we have discrete field inside it so we have a list of all distinct values and we can show this filter on the right
side if you just right click on the subcategory over here and then select show filter so now we have it on the
right side and we can now include or exclude values and now let's see what's going to happen if we put in the filters
continuous fi so let's take the sales again since it's continuous fi but instead of taking it from the left side
here from datab ban you can take it from the shields by holding alt and then drag and drop in the filters so since it's
continuous field and a measure tblo going to ask us first do we want to do the filter in all values or after we do
the calculations so let's go with the sum over here since we have it as a sum so I'm just going to click on the sum
and go next and this is exactly what going to happen if you have continuous field as a filter you will get to range
it has a start and ends so you don't have like distinct values of all the sales you will get a range of values and
you have to define the start and the end and here we have different options about the range but we're going to stay with
the first one so let's hit okay and now I want to show the filter on the right side so let's go over here right click
on show filter and now on the right side you can see exactly the difference between discrete and continuous fields
in filters so let me just extend it over here you see the sales is continuous and we have range so we can filter like this
by changing the start and the end of the range but with the discrete filter we have all members of the field and we can
decide on each value individually so we can just select and deselect those values all right so now let's move to
the next point we're going to talk about the changes in the view discrete Fields create the headers of the visualizations
where the continuous Fields creates the axis of visualizations okay so now let's see what this means in our view as you
can see the subcategory is a discrete field and the sales is continuous field and in this view over here we have three
things we have the marks those parts and on the left side we have the subcategory and we call those informations as
headers and the third information we have the access of the view so what is the difference between headers and axes
the discrete Fields like subcategory always create the header of the view and in the header over here you have like
list of all distinct values inside our data set exactly as it is but The Continuous field like the sales create
the axis of the visualization and it's like the values inside a filter it's a range that has a starts and ends and
unlike the headers you cannot see in the AIS all the possible values individually so you have a range with start and ends
and in between we have pens so discrete Fields create the headers and continuous Fields create the
axis all right so the next point we're going to talk about sorting data in discrete fields we have many options in
order to sort the data but with the continuous fields in Tableau it is very limited so let's see an example so we're
going to stay with the same example and we going to start with the discrete field subcategory so in order to the
data in the discret field just right click on the subcategory over here on the shelf or you can go to the header
it's exactly the same so right click on the subcategory and then we can select over here the sort so select that and
now we have extra window to set up the sort so as you can see here we have many different options like alphabetic Field
Manual and so on so let's go with the manual over here and here again since subcategory is discrete Fields we're
going to get the list of all distinct values and then we can change the order for example by just clicking on the
applications we just can break it down and we can take the storage and bring it up blenders down and so on so we can do
it manually without any rule so as you can see as I'm changing the values their order in the visualization is as well
changing so if you want to sort the data we're going to use the discrete fields in order to do that since we have many
options and now let's check the continuous Fields so I'm going to close this so now if you go to the continuous
fields on the sales right click on it we don't have here an option to sort the data like in the discrete Fields but
instead we have only one option if you hover on the sales we have this very small icon and we can use it in order
order to sort the data ascending or descending so just click on that and as you can see now the data is sorted by
descending values and if you click on that again you will get the data as ascending so sorting the data using
continuous field is very limited but instead of that we can use the discrete fields in order to sort the data since
we have many options okay so now let's move to the next one and this is really important to
understand what is really the purpose of having continuous and discrete in Tableau the main use case of using the
discrete values is to do a deep Dives analyzis in specific scenario and in the other hand we're going to use the
continuous values to see the big picture and do Trend analyzes let's have an example now we're going to create a new
View using the big data source since we have more data and we're going to go to the table orders let's take the order
date just drag and drop it on the columns and then we're going to take one measure let's say the quantity drag and
drop it on the rows and now as you can see the order date is a discrete field and we have 5 years of that down but now
what we're going to do we're going to go to the order date right click on it and we want to see more details so just go
to the exact date over here and now as you can see Tableau did convert it automatically from discrete to
continuous value and we have it as a green pill and that's because we have a lot of order dates and Tableau try to
bring it all in one picture and you can see now the order date created an axis with a range of dates so having
continuous Fields you have all the data in one big picture and that's going to help you to find any Trend in your data
so now let's go and convert the order date to a discrete field so in order to do that we're going to go to the order
date right click on it and click on discrete as you can see now we just broke the chain and we broke the
visualizations into individual dates and now because of that we have the header and we have all the distinct values
inside our data so we have all the days all the months of the five years in one visual so that having the order day as a
discret we cannot really do any Trend analyzis over here because it's really huge visualization so after we converted
the order date from continuous to discret we lost the big picture and now it's really hard to do any Trend
analysis but now instead of doing Trend analyzis we can do now a deep dive details analyzes for each individual
dates in order to analyze a specific problem or scenario or to answer the question why do we have in the first
place a trend so you can check the value of each date individually and we usually use the bar visualizations for the
discret and the line visualizations for the continuous so let's change that I will go over here on the marks and
instead of automatic I will move it to bar so we have it now here as a bar and I'm
going to just duplicate this sheet and bring the order date as a continuous and then change the
visualizations to automatic and now I just moved both of the views into One dashboard in order to see the
differences between continuous and discret so as you can see with the continuous if you want to make like
Trend analyzes seeing the big picture or you going to make like a report for the management without showing a lot of
details then go and use the continuous field and now if you look at the visualizations with the discrete fields
you can use that if the task or the requirement is to do deep dive analysis in the data and evaluate each data
individually so the main purpose of having discret is to do detailed analyzes where the purpose of continuous
values is to do Trend analyzes all right so now let's have a summary where we're going to compare
both of the discrete and continuous side byid in order to understand the differences between them let's start
with the definitions discrete values are disconnected separated Val values and continuous values are connected unbroken
chain of values for example in discret between 0 and 10 we have finite number of values we have exactly 11 values and
in continuous between one and two we have infinite number of values next one is about the colors discrete fields are
the blue pills and continuous fields are the green pills moving on to filters discrete Fields generate filters with a
distinct list of all values available in the data set and in the other hand the prous Fields generate a range filter
that has start and end values and next point is about the views discrete Fields can generate the header of the view
showing all possible values and the continuous Fields generate the axis of the view again it's like range of values
then we have sorting you can use discrete fields to sort your data using different options but if you sort your
data using continuous Fields you're going to have very limited options we have only ascending or descending and
finally we're going to talk about thep purposes the main purpose of the discrete is to analyze a specific
scenario like you are doing a deep dive analysis in a specific issue but the main purpose of the continuous is to
understand the big picture from the data in order to do for example Trend analyzes of your data so these are the
main differences between discret and continuous Fields all right guys so now what I'm
going to show you is how those different metadata Concepts like data types dimensions and measures discret and
continuous are related to each other all right so now we have a field in our data and in Tableau we can assign it to
different data types so it could be string or poon with true and false or a date and we have as well date and time
or a number whether it's whole or decimal and now next TBL can assign it to another metadata info either
Dimension or measure any data type that is not a number it's going to be Dimension so string buan and date all of
them going to be automatically dimension cannot convert it to a measure and if the data type is number we could have it
as a measure or Dimension if it makes sense to do aggregation and next T going to assign this field to the third
metadata concept discrete or continuous if we have a dimension field with a data type string it could be only discrete we
cannot convert it to continuous like in our data set we have the category the first name the country all those fields
are string Dimension and discret you cannot change it to anything else the same goes for the data type buan it
could be only Dimension and and only discret but now if we have a dimension filled with a data type date or date
time as you saw in our examples it could be continuous or discret we can have both and now to the last one if we have
a field with a data type number it doesn't matter whether it's Dimension or measure we can have this field as
continuous and as well as discrete all right guys so with this you have big picture for all those confusing Concepts
in metadata in Tableau all right everyone so we have now better understanding about the data types and
roles in Tableau and these important Concepts and in the next section we will learn about renaming and aliases in
Tableau how to rename things in Tableau as we are preparing our data sources what we usually do is that we're going
to go and rename stuff like renaming tables columns and even give ilas to our data so first I'm going to introduce you
to the different naming conventions that each developer should know and after that you're going to learn the different
techniques on how to rename fields and tables in Tableau and at the end you're going to learn the different method on
how to add aliases to your data in Tableau so let's start first by learning the different naming conventions and
what are the differences between them so now let's go sometimes in real life projects the
source of your data might contain technical or unfriendly names and when you are creating visualizations for the
users or your colleagues you have to make sure that you are using friendly names that are easy to understand and to
read and that's why after you connect your data to Tableau Data sources Tableau will start cleaning up and
renaming the fields and the tables to more friendly format and the format is following specific naming convention
that is decided from the Tableau team which is really great so let's understand first what is naming
convention naming conventions are set of rules and guidelines that could be used in order to give names for things like
tables Fields functions and variables in consistent and understandable way let's say for example we have the two words
hello word in order to create a naming convention we have to decide in two things first the word itself how we
going to write it here we have three ways we can use the lower case or we can decide to go with the uppercase or we
could use the capital letters and the second thing to decide is the separator between words so between hello and word
we have here white space here we have different options you could use dots uncore FLH Whit space or even nothing so
now for example let's say we're going to go with the lowercase and the separator underscore then we're going to have the
following name hellcore world so with that we have a naming convention that we're going to follow through all the
projects and it's really easy to follow and at the same time it's very important to decide on the naming convention for
your data model especially at the start of your project and if you don't do that I promise you the look and feeling of
your visualizations and dashboard going to look really bad and the whole project going to look unprofessional and
inconsistent and one more thing project team decides on different naming conventions so there is no really right
and wrong here all right everyone so now I'm going to walk you through the most common
naming conventions used in programming languages the first naming convention is the snake
case k is going to use the lower case in all the words and going to separate them using the underscore so the name at the
end is going to look like snake all right so our example going to be the customer name and we're going to work
with this table to fill all the different naming conventions an example of the output the rules for the lit case
and the separators and in which applications and programming languages we can find this rule where we're going
to start with the snake case the letter case going to be here lower case and the separator going to be the underscore so
if we follow those rules with the example we're going to have a lowercase customer
underscore name and we can find those formats in Python PHP and Ruby so the snake format is really easy and popular
and you can find it like almost everywhere and now we're going to talk about the next naming convention we have
the camel case and here we have another naming convention that looks like an animal so
in the camel case only the first word going to be lowercase but then all the following words going to be capitalized
and between the words there is nothing no separators no dots underscores dashes or anything so at the end we're going to
have the shape of camel all right so that means we have the second naming convention we have the camel case the
rule for the letter K is going to be the following the first words going to be lower and the rest of the words going to
be capitalized for the second rule we have the separation there is no separation there is nothing between the
words so here we're going to write no separation so now if you apply those two rules in our example the customer name
we're going to have the following output so the first one going to be everything lowercase
customer there is no separations that means we're going to start immediately with the second word but the second word
going to be capitalized so it's going to be name like this and we can see that Camel case is widely used in programming
languages like Java JavaScript and typescripts okay so that means we have the third naming convention we have the
the Pascal case it's very similar to the camel case so the rule says all the words going to be capitalized so here we
have capitalized and the separations there is no separation like the camel case so there is nothing so if you
follow those two rules on the customer name we're going to have the following outputs so the first word going to be
customer capitalized no oparation then a capitalized name and we can find this naming convention the Pascal case is
used in programming languages like Java and c I like this naming convention I used it in many
projects all right the next naming convention going to be the Kebab case and I think by now the one who
named those naming conventions should be an Arabic dude as you can see we have all the words are lower cased in the
skewer and separated with dashes so the name going to look like a delicious hot kebab skewer so now the fourth one we
have the Kebab case and the rule going to say okay the L case going to be lowercased like the snake case and the
separation going to be here the dash so if we follow those two rules on the customer name in our example we going to
have the foll output it's really easy going to be customer or lower then a dash then name and if you are web
developer or designer I think you know about this naming convention because it is widely used in HTML and CSS I think
it's like the Snak casee it's really easy to follow and now we have another naming
convention this one is very important and we call it a title case it has nothing to do with animals or Foods
sadly so we have here title case they're all going to say okay the word's going to be capitalized and we're going to
separate the words with a white space so here we're going to have space so now if you follow those two rules in our
example we're going to have capitalized customer then space then capitalized name like this so why it's important
because this one is the naming convention that Tableau team did decide to go with so you can see this naming
convention in Tableau so Tableau currently is enforcing this naming convention in all your data so once you
connect your data to Tableau Tableau going to clean up and rename everything following this rule well if you look at
