Anupam Baral/blog
AI & Machine Learning

AI vs Machine Learning: A Beginner’s Guide to EDA, Data Cleaning & the Complete ML Workflow

Artificial Intelligence (AI) and Machine Learning (ML) are everywhere today. They recommend movies on Netflix, filter spam emails, help doctors detect diseases, power self-driving ...

Anupam BaralAnupam Baral
··9 min read
AI vs Machine Learning: A Beginner’s Guide to EDA, Data Cleaning & the Complete ML Workflow

From Raw Data to Smart Predictions: A Beginner’s Guide to AI, Machine Learning, and Data Preparation

fig : Machine Learning workflow.
fig : Machine Learning workflow.

Introduction

Artificial Intelligence (AI) and Machine Learning (ML) are everywhere today. They recommend movies on Netflix, filter spam emails, help doctors detect diseases, power self-driving cars, and even generate text and images.

If you’re just starting your AI or ML journey, it’s easy to feel overwhelmed by all the new terms. You might hear words like EDA , feature engineering , preprocessing , or Random Forest , and wonder where they all fit together.

The good news is that machine learning isn’t just about building models. In fact, most of the work happens before training a model.

This guide explains the complete machine learning workflow in simple language, from collecting data to training a model, with practical examples along the way.

What is Artificial Intelligence (AI)?

Artificial Intelligence (AI) is the broad field of creating computers that can perform tasks requiring human intelligence.

These tasks include:

  • Recognizing images

  • Understanding language

  • Playing games

  • Making recommendations

  • Solving problems

  • Learning from experience

Think of AI as the big umbrella .

Examples:

  • ChatGPT

  • Google Assistant

  • Face Unlock on smartphones

  • Self-driving cars

  • AI chatbots

What is Machine Learning (ML)?

Machine Learning is a subset of AI .

Instead of programming every rule manually, we give the computer lots of examples (data), and it learns patterns by itself.

For example:

Instead of writing hundreds of rules to identify spam emails, we give the model thousands of spam and non-spam emails.

The model learns the difference automatically.

Simple idea:

Data → Learn Patterns → Make Predictions

Examples:

  • Predict house prices

  • Detect fraud

  • Recommend products

  • Predict customer churn

  • Classify images

AI vs Machine Learning

fig : AI vs Machine Learning
fig : AI vs Machine Learning

Artificial Intelligence (AI)

- Broad field of computer science - Focuses on creating intelligent systems - May use rules, logic, machine learning, or other techniques - Goal: Perform tasks that normally require human intelligence - Example: ChatGPT, Siri, self-driving cars

Machine Learning (ML)

- A subset of Artificial Intelligence - Focuses on learning patterns from data - Learns automatically from examples instead of explicit rules - Goal: Make predictions or decisions based on data - Example: Spam detection, Netflix recommendations, house price prediction

A simple analogy:

  • AI = Entire hospital

  • ML = One doctor inside the hospital

Machine Learning is one important way to build AI systems, but it isn’t the only way.

The Machine Learning Workflow

Many beginners think machine learning starts with selecting an algorithm.

Actually, it starts much earlier.

A typical workflow looks like this:

  1. Problem Classification

  2. Collect data

  3. Perform Exploratory Data Analysis (EDA)

  4. Clean the data

  5. Preprocess the data

  6. Extract or engineer features

  7. Train a model

  8. Evaluate the model

  9. Improve and deploy

Each step is important because poor data usually leads to poor predictions.

Step 1: Data Collection

Everything begins with data.

Data can come from:

  • CSV files

  • Excel spreadsheets

  • Databases

  • APIs

  • Sensors

  • User surveys

  • Websites

  • Business applications

Example:

Suppose we want to predict house prices.

Our dataset might contain:

  • Number of bedrooms

  • House size

  • Location

  • Year built

  • Garage

  • Price

Without good data, no machine learning model can perform well.

Step 2: What is Exploratory Data Analysis (EDA)?

fig : Exploratory Data Analysis (EDA)
fig : Exploratory Data Analysis (EDA)

EDA stands for Exploratory Data Analysis .

Before building any model, we need to understand the data.

Think of EDA as becoming familiar with your dataset before asking it to solve a problem.

During EDA, we ask questions like:

  • How many rows and columns are there?

  • What features are available?

  • Are values missing?

  • Are there unusual values?

  • What patterns exist?

  • Are variables related?

EDA helps uncover problems early and guides later decisions.

Common Steps in EDA

1. Understand the Dataset

Look at:

  • Number of rows

  • Number of columns

  • Data types

  • Feature names

2. View Sample Data

Inspect the first few rows to understand the structure.

Questions to ask:

  • Does everything look correct?

  • Are columns meaningful?

3. Check Missing Values

Some records may be incomplete.

Example:

AgeSalary25$45,00030Missing27$52,000

Missing values need attention before training.

4. Explore Distributions

Look at how values are spread.

Examples:

  • Age distribution

  • Salary distribution

  • Product prices

This helps identify unusual patterns.

5. Detect Outliers

Outliers are values that look very different from the rest.

Example:

Most salaries:

  • $40,000

  • $42,000

  • $45,000

One salary:

  • $2,000,000

That value deserves investigation.

6. Study Relationships

Questions include:

  • Does experience affect salary?

  • Does house size affect price?

  • Does age influence spending?

Understanding relationships helps with feature selection later.

After EDA: Data Cleaning

Once we understand the dataset, we begin cleaning it.

This is one of the most important stages in machine learning.

Many practitioners estimate that around 80% of the work in a machine learning project involves preparing and cleaning data , rather than training models.

A great algorithm cannot compensate for poor-quality data.

Common Data Cleaning Tasks

Handle Missing Values

Missing information is very common.

Options include:

  • Remove incomplete rows

  • Fill with the average

  • Fill with the median

  • Fill with the most common value

  • Predict missing values using other features

Remove Duplicate Data

Duplicate records can bias the model.

Always check for repeated rows.

Correct Errors

Examples:

Instead of:

  • Kathmanduu

  • kathmandu

  • KTM

Standardize them to:

  • Kathmandu

Consistency matters.

Fix Data Types

Examples:

Age stored as text:

“25”

should become:

25

Incorrect data types often cause errors later.

Handle Outliers

Not every outlier is incorrect.

Sometimes it represents a real event.

Decide whether to:

  • Keep it

  • Remove it

  • Cap extreme values

This decision should be based on context, not automatically.

Why Data Cleaning Is Difficult

Cleaning data is rarely straightforward.

Common challenges include:

  • Missing information

  • Duplicate records

  • Typing mistakes

  • Inconsistent formats

  • Mixed units

  • Invalid values

  • Incorrect labels

  • Unexpected categories

  • Data collected from multiple sources

  • Human entry errors

Real-world datasets are often messy.

That is why data cleaning consumes such a large portion of a machine learning project.

Data Preprocessing

After cleaning, we prepare the data so machine learning algorithms can use it effectively.

This stage is called data preprocessing .

Common preprocessing tasks include:

Encode Categories

Algorithms understand numbers, not text.

Example:

Color: Red Blue Green

becomes numerical values using encoding techniques.

Scale Numerical Features

Features with very different ranges can affect some algorithms.

Example:

  • Salary: 80,000

  • Age: 25

Scaling brings them into comparable ranges.

Split the Dataset

Typically:

  • Training data

  • Validation data (optional)

  • Testing data

The model learns from the training set and is evaluated on unseen data.

Feature Extraction and Feature Engineering

A feature is an input used by a machine learning model.

Sometimes the original data isn’t the most useful.

We create better features from existing ones.

Example:

Instead of storing:

  • Birth year

Create:

  • Current age

Another example:

Instead of:

  • Purchase date

Create:

  • Day of the week

  • Month

  • Holiday indicator

Better features often improve model performance more than switching algorithms.

Types of Machine Learning

1. Supervised Learning

fig : Supervised Learning
fig : Supervised Learning

The data includes correct answers (labels).

Goal:

Learn to predict those answers.

Examples:

  • Predict house prices

  • Detect spam

  • Predict loan approval

Common models:

  • Linear Regression

  • Logistic Regression

  • Decision Tree

  • Random Forest

  • Support Vector Machine (SVM)

2. Unsupervised Learning

fig : Unsupervised Learning
fig : Unsupervised Learning

The data has no labels.

The algorithm discovers hidden patterns.

Examples:

  • Customer segmentation

  • Product grouping

  • Finding unusual behavior

Common models:

  • K-Means Clustering

  • DBSCAN

  • Hierarchical Clustering

  • PCA (for dimensionality reduction)

3. Reinforcement Learning

fig : Reinforcement Learning
fig : Reinforcement Learning

An agent learns through rewards and penalties.

The goal is to maximize long-term rewards.

Examples:

  • Game-playing AI

  • Robotics

  • Self-driving systems

  • Resource optimization

Common algorithms:

  • Q-Learning

  • Deep Q Networks (DQN)

  • Policy Gradient methods

A Simple End-to-End Example

Imagine you want to predict whether a student will pass an exam.

Step 1: Collect Data

Gather:

  • Study hours

  • Attendance

  • Previous grades

  • Sleep hours

  • Final result (Pass/Fail)

Step 2: Perform EDA

You notice:

  • Some attendance values are missing.

  • A few study-hour entries are impossible (such as 200 hours in one week).

  • Students with higher attendance often perform better.

Step 3: Clean the Data

  • Fill missing attendance values.

  • Remove clearly incorrect entries.

  • Correct inconsistent labels like “pass”, “Pass”, and “PASS” so they all match.

Step 4: Preprocess

  • Convert Pass/Fail into numerical labels.

  • Scale numerical features if needed.

  • Split the dataset into training and testing sets.

Step 5: Feature Engineering

Create a new feature:

Study Efficiency = Study Hours ÷ Attendance

This may provide more insight than either value alone.

Step 6: Train a Model

Choose a simple supervised learning algorithm, such as a Decision Tree or Logistic Regression.

Step 7: Evaluate

Test the model on unseen students.

If the predictions are accurate enough, the model can be improved further or deployed.

Tips for Writing Beginner-Friendly AI and ML Blogs

If you’re writing for newcomers, focus on understanding rather than complexity.

Here are a few tips:

  • Start with a real-world problem readers recognize.

  • Explain one idea at a time.

  • Use simple analogies.

  • Avoid unnecessary mathematical notation.

  • Include diagrams or workflow illustrations where possible.

  • Use short paragraphs and descriptive headings.

  • Define technical terms the first time you mention them.

  • End each section with a short takeaway.

  • Encourage readers to experiment with small datasets before moving to larger projects.

Remember, readers often stay engaged because the concepts feel approachable, not because the content is packed with jargon.

Quick Glossary

  • AI: Systems designed to perform tasks that normally require human intelligence.

  • Machine Learning: A branch of AI where computers learn patterns from data.

  • Dataset: A collection of data used for analysis or training.

  • EDA: Exploring data to understand its quality, structure, and patterns.

  • Feature: An input variable used by a machine learning model.

  • Model: The algorithm that learns from data to make predictions.

  • Training: Teaching a model using historical data.

  • Prediction: The output generated by a trained model.

Beginner’s Checklist

Before training a model, ask yourself:

  • Have I collected enough relevant data?

  • Have I explored the dataset with EDA?

  • Have I handled missing values?

  • Have I removed duplicates and corrected obvious errors?

  • Have I preprocessed the data correctly?

  • Have I created useful features?

  • Have I split the data into training and testing sets?

  • Am I using the right type of machine learning for my problem?

  • Have I evaluated the model on unseen data?

If you can answer “yes” to each of these questions, you’re building on a strong foundation.

fig : AI / ML summery
fig : AI / ML summery

Conclusion

Machine learning is much more than choosing an algorithm. The journey begins with collecting reliable data, understanding it through Exploratory Data Analysis, cleaning inconsistencies, preprocessing features, and only then training a model.

Many successful machine learning projects aren’t won by using the most advanced algorithm — they succeed because the data is well understood and carefully prepared. By mastering these fundamentals first, you’ll build models that are more accurate, more reliable, and easier to improve over time.

As you continue learning, start with small datasets, practice each step of the workflow, and remember that becoming skilled at working with data is just as valuable as learning new algorithms. Strong foundations in data preparation will serve you well no matter which area of AI or ML you explore next.