Python Data Preprocessing
July 12, 2024 · 1 min read
If you have any questions, feel free to comment below. Click the block can copy the code.
And if you think it's helpful to you, just click on the ads which can support this site. Thanks!
Reading the Dataset #
The data is stored in the CSV (comma-separated values) file ../data/house_tiny.csv.
import os
os.makedirs(os.path.join('..', 'data'), exist_ok=True)
data_file = os.path.join('..', 'data', 'house_tiny.csv')
with open(data_file, 'w') as f:
f.write('NumRooms,Alley,Price\n') # Column names
f.write('NA,Pave,127500\n') # Each row represents one data sample
f.write('2,NA,106000\n')
f.write('4,NA,178100\n')
f.write('NA,NA,140000\n')
# To load the raw dataset from the CSV file we created, import `pandas` and call `read_csv`.
# !pip install pandas
import pandas as pd
data = pd.read_csv(data_file)
print(data)
# NumRooms Alley Price
# 0 NaN Pave 127500
# 1 2.0 NaN 106000
# 2 4.0 NaN 178100
# 3 NaN NaN 140000
Handling Missing Values #
Typical ways to handle missing data include:
- Imputation (filling in missing values with replacement values)
- Deletion (simply ignoring missing values)
Using the positional indexer iloc, we replace missing numerical values in inputs (the “NaN” entries) with the mean of their column.
inputs, outputs = data.iloc[:, 0:2], data.iloc[:, 2]
inputs = inputs.fillna(inputs.mean()) # Fill with the mean of the same column
print(inputs)
NumRooms Alley
0 3.0 Pave
1 2.0 NaN
2 4.0 NaN
3 3.0 NaN
# `pandas` can automatically convert this column into two columns: Alley_Pave and Alley_nan
inputs = pd.get_dummies(inputs, dummy_na=True)
print(inputs)
NumRooms Alley_Pave Alley_nan
0 3.0 1 0
1 2.0 0 1
2 4.0 0 1
3 3.0 0 1
Converting to Tensor Format #
Now that all entries in inputs and outputs are numerical, they can be converted to tensors.
import torch
X = torch.tensor(inputs.to_numpy(dtype=float))
y = torch.tensor(outputs.to_numpy(dtype=float))
X, y
(tensor([[3., 1., 0.],
[2., 0., 1.],
[4., 0., 1.],
[3., 0., 1.]], dtype=torch.float64),
tensor([127500., 106000., 178100., 140000.], dtype=torch.float64))
Summary #
- Write with
f.write()
with open(data_file, 'w') as f:
f.write('NumRooms,Alley,Price\n') # Column names
- Load with
pd.read_csv()
import pandas as pd
data = pd.read_csv(data_file)
- Positional indexing with
.iloc()
inputs, outputs = data.iloc[:, 0:2], data.iloc[:, 2]
- Fill with the mean using
.fillna()
inputs = inputs.fillna(inputs.mean())
- Convert NaN values to indicator columns
pd.get_dummies(inputs, dummy_na=True)
- Convert numerical values to tensors with
torch.tensor(a.to_numpy(dtype=float))
torch.tensor(outputs.to_numpy(dtype=float))
Related readings
- Python Data Manipulation
- Density Plotting
- Types of Univariate Plots for Continuous Variables
- SciencePlots Basics and Features
- ProPlot Basics and Features
If you want to follow my updates, or have a coffee chat with me, feel free to connect with me: