AI vs Machine Learning: A Beginner’s Guide to EDA, Data Cleaning & the Complete ML Workflow
Artificial Intelligence (AI) and Machine Learning (ML) are everywhere today. They recommend movies on Netflix, filter spam emails, help doctors detect diseases, power self-driving ...

From Raw Data to Smart Predictions: A Beginner’s Guide to AI, Machine Learning, and Data Preparation

Introduction
Artificial Intelligence (AI) and Machine Learning (ML) are everywhere today. They recommend movies on Netflix, filter spam emails, help doctors detect diseases, power self-driving cars, and even generate text and images.
If you’re just starting your AI or ML journey, it’s easy to feel overwhelmed by all the new terms. You might hear words like EDA , feature engineering , preprocessing , or Random Forest , and wonder where they all fit together.
The good news is that machine learning isn’t just about building models. In fact, most of the work happens before training a model.
This guide explains the complete machine learning workflow in simple language, from collecting data to training a model, with practical examples along the way.
What is Artificial Intelligence (AI)?
Artificial Intelligence (AI) is the broad field of creating computers that can perform tasks requiring human intelligence.
These tasks include:
Recognizing images
Understanding language
Playing games
Making recommendations
Solving problems
Learning from experience
Think of AI as the big umbrella .
Examples:
ChatGPT
Google Assistant
Face Unlock on smartphones
Self-driving cars
AI chatbots
What is Machine Learning (ML)?
Machine Learning is a subset of AI .
Instead of programming every rule manually, we give the computer lots of examples (data), and it learns patterns by itself.
For example:
Instead of writing hundreds of rules to identify spam emails, we give the model thousands of spam and non-spam emails.
The model learns the difference automatically.
Simple idea:
Data → Learn Patterns → Make Predictions
Examples:
Predict house prices
Detect fraud
Recommend products
Predict customer churn
Classify images
AI vs Machine Learning

Artificial Intelligence (AI)
- Broad field of computer science - Focuses on creating intelligent systems - May use rules, logic, machine learning, or other techniques - Goal: Perform tasks that normally require human intelligence - Example: ChatGPT, Siri, self-driving cars
Machine Learning (ML)
- A subset of Artificial Intelligence - Focuses on learning patterns from data - Learns automatically from examples instead of explicit rules - Goal: Make predictions or decisions based on data - Example: Spam detection, Netflix recommendations, house price prediction
A simple analogy:
AI = Entire hospital
ML = One doctor inside the hospital
Machine Learning is one important way to build AI systems, but it isn’t the only way.
The Machine Learning Workflow
Many beginners think machine learning starts with selecting an algorithm.
Actually, it starts much earlier.
A typical workflow looks like this:
Problem Classification
Collect data
Perform Exploratory Data Analysis (EDA)
Clean the data
Preprocess the data
Extract or engineer features
Train a model
Evaluate the model
Improve and deploy
Each step is important because poor data usually leads to poor predictions.
Step 1: Data Collection
Everything begins with data.
Data can come from:
CSV files
Excel spreadsheets
Databases
APIs
Sensors
User surveys
Websites
Business applications
Example:
Suppose we want to predict house prices.
Our dataset might contain:
Number of bedrooms
House size
Location
Year built
Garage
Price
Without good data, no machine learning model can perform well.
Step 2: What is Exploratory Data Analysis (EDA)?

EDA stands for Exploratory Data Analysis .
Before building any model, we need to understand the data.
Think of EDA as becoming familiar with your dataset before asking it to solve a problem.
During EDA, we ask questions like:
How many rows and columns are there?
What features are available?
Are values missing?
Are there unusual values?
What patterns exist?
Are variables related?
EDA helps uncover problems early and guides later decisions.
Common Steps in EDA
1. Understand the Dataset
Look at:
Number of rows
Number of columns
Data types
Feature names
2. View Sample Data
Inspect the first few rows to understand the structure.
Questions to ask:
Does everything look correct?
Are columns meaningful?
3. Check Missing Values
Some records may be incomplete.
Example:
AgeSalary25$45,00030Missing27$52,000
Missing values need attention before training.
4. Explore Distributions
Look at how values are spread.
Examples:
Age distribution
Salary distribution
Product prices
This helps identify unusual patterns.
5. Detect Outliers
Outliers are values that look very different from the rest.
Example:
Most salaries:
$40,000
$42,000
$45,000
One salary:
$2,000,000
That value deserves investigation.
6. Study Relationships
Questions include:
Does experience affect salary?
Does house size affect price?
Does age influence spending?
Understanding relationships helps with feature selection later.
After EDA: Data Cleaning
Once we understand the dataset, we begin cleaning it.
This is one of the most important stages in machine learning.
Many practitioners estimate that around 80% of the work in a machine learning project involves preparing and cleaning data , rather than training models.
A great algorithm cannot compensate for poor-quality data.
Common Data Cleaning Tasks
Handle Missing Values
Missing information is very common.
Options include:
Remove incomplete rows
Fill with the average
Fill with the median
Fill with the most common value
Predict missing values using other features
Remove Duplicate Data
Duplicate records can bias the model.
Always check for repeated rows.
Correct Errors
Examples:
Instead of:
Kathmanduu
kathmandu
KTM
Standardize them to:
Kathmandu
Consistency matters.
Fix Data Types
Examples:
Age stored as text:
“25”
should become:
25
Incorrect data types often cause errors later.
Handle Outliers
Not every outlier is incorrect.
Sometimes it represents a real event.
Decide whether to:
Keep it
Remove it
Cap extreme values
This decision should be based on context, not automatically.
Why Data Cleaning Is Difficult
Cleaning data is rarely straightforward.
Common challenges include:
Missing information
Duplicate records
Typing mistakes
Inconsistent formats
Mixed units
Invalid values
Incorrect labels
Unexpected categories
Data collected from multiple sources
Human entry errors
Real-world datasets are often messy.
That is why data cleaning consumes such a large portion of a machine learning project.
Data Preprocessing
After cleaning, we prepare the data so machine learning algorithms can use it effectively.
This stage is called data preprocessing .
Common preprocessing tasks include:
Encode Categories
Algorithms understand numbers, not text.
Example:
Color: Red Blue Green
becomes numerical values using encoding techniques.
Scale Numerical Features
Features with very different ranges can affect some algorithms.
Example:
Salary: 80,000
Age: 25
Scaling brings them into comparable ranges.
Split the Dataset
Typically:
Training data
Validation data (optional)
Testing data
The model learns from the training set and is evaluated on unseen data.
Feature Extraction and Feature Engineering
A feature is an input used by a machine learning model.
Sometimes the original data isn’t the most useful.
We create better features from existing ones.
Example:
Instead of storing:
Birth year
Create:
Current age
Another example:
Instead of:
Purchase date
Create:
Day of the week
Month
Holiday indicator
Better features often improve model performance more than switching algorithms.
Types of Machine Learning
1. Supervised Learning

The data includes correct answers (labels).
Goal:
Learn to predict those answers.
Examples:
Predict house prices
Detect spam
Predict loan approval
Common models:
Linear Regression
Logistic Regression
Decision Tree
Random Forest
Support Vector Machine (SVM)
2. Unsupervised Learning

The data has no labels.
The algorithm discovers hidden patterns.
Examples:
Customer segmentation
Product grouping
Finding unusual behavior
Common models:
K-Means Clustering
DBSCAN
Hierarchical Clustering
PCA (for dimensionality reduction)
3. Reinforcement Learning

An agent learns through rewards and penalties.
The goal is to maximize long-term rewards.
Examples:
Game-playing AI
Robotics
Self-driving systems
Resource optimization
Common algorithms:
Q-Learning
Deep Q Networks (DQN)
Policy Gradient methods
A Simple End-to-End Example
Imagine you want to predict whether a student will pass an exam.
Step 1: Collect Data
Gather:
Study hours
Attendance
Previous grades
Sleep hours
Final result (Pass/Fail)
Step 2: Perform EDA
You notice:
Some attendance values are missing.
A few study-hour entries are impossible (such as 200 hours in one week).
Students with higher attendance often perform better.
Step 3: Clean the Data
Fill missing attendance values.
Remove clearly incorrect entries.
Correct inconsistent labels like “pass”, “Pass”, and “PASS” so they all match.
Step 4: Preprocess
Convert Pass/Fail into numerical labels.
Scale numerical features if needed.
Split the dataset into training and testing sets.
Step 5: Feature Engineering
Create a new feature:
Study Efficiency = Study Hours ÷ Attendance
This may provide more insight than either value alone.
Step 6: Train a Model
Choose a simple supervised learning algorithm, such as a Decision Tree or Logistic Regression.
Step 7: Evaluate
Test the model on unseen students.
If the predictions are accurate enough, the model can be improved further or deployed.
Tips for Writing Beginner-Friendly AI and ML Blogs
If you’re writing for newcomers, focus on understanding rather than complexity.
Here are a few tips:
Start with a real-world problem readers recognize.
Explain one idea at a time.
Use simple analogies.
Avoid unnecessary mathematical notation.
Include diagrams or workflow illustrations where possible.
Use short paragraphs and descriptive headings.
Define technical terms the first time you mention them.
End each section with a short takeaway.
Encourage readers to experiment with small datasets before moving to larger projects.
Remember, readers often stay engaged because the concepts feel approachable, not because the content is packed with jargon.
Quick Glossary
AI: Systems designed to perform tasks that normally require human intelligence.
Machine Learning: A branch of AI where computers learn patterns from data.
Dataset: A collection of data used for analysis or training.
EDA: Exploring data to understand its quality, structure, and patterns.
Feature: An input variable used by a machine learning model.
Model: The algorithm that learns from data to make predictions.
Training: Teaching a model using historical data.
Prediction: The output generated by a trained model.
Beginner’s Checklist
Before training a model, ask yourself:
Have I collected enough relevant data?
Have I explored the dataset with EDA?
Have I handled missing values?
Have I removed duplicates and corrected obvious errors?
Have I preprocessed the data correctly?
Have I created useful features?
Have I split the data into training and testing sets?
Am I using the right type of machine learning for my problem?
Have I evaluated the model on unseen data?
If you can answer “yes” to each of these questions, you’re building on a strong foundation.

Conclusion
Machine learning is much more than choosing an algorithm. The journey begins with collecting reliable data, understanding it through Exploratory Data Analysis, cleaning inconsistencies, preprocessing features, and only then training a model.
Many successful machine learning projects aren’t won by using the most advanced algorithm — they succeed because the data is well understood and carefully prepared. By mastering these fundamentals first, you’ll build models that are more accurate, more reliable, and easier to improve over time.
As you continue learning, start with small datasets, practice each step of the workflow, and remember that becoming skilled at working with data is just as valuable as learning new algorithms. Strong foundations in data preparation will serve you well no matter which area of AI or ML you explore next.