timerring

Python Data Preprocessing

July 12, 2024 · 1 min read
Tutorial
Python
If you have any questions, feel free to comment below. Click the block can copy the code.
And if you think it's helpful to you, just click on the ads which can support this site. Thanks!

Reading the Dataset #

The data is stored in the CSV (comma-separated values) file ../data/house_tiny.csv.

import os

os.makedirs(os.path.join('..', 'data'), exist_ok=True)
data_file = os.path.join('..', 'data', 'house_tiny.csv')
with open(data_file, 'w') as f:
    f.write('NumRooms,Alley,Price\n')  # Column names
    f.write('NA,Pave,127500\n')  # Each row represents one data sample
    f.write('2,NA,106000\n')
    f.write('4,NA,178100\n')
    f.write('NA,NA,140000\n')

# To load the raw dataset from the CSV file we created, import `pandas` and call `read_csv`.
# !pip install pandas
import pandas as pd

data = pd.read_csv(data_file)
print(data)

#    NumRooms Alley   Price
# 0       NaN  Pave  127500
# 1       2.0   NaN  106000
# 2       4.0   NaN  178100
# 3       NaN   NaN  140000

Handling Missing Values #

Typical ways to handle missing data include:

  • Imputation (filling in missing values with replacement values)
  • Deletion (simply ignoring missing values)

Using the positional indexer iloc, we replace missing numerical values in inputs (the “NaN” entries) with the mean of their column.

inputs, outputs = data.iloc[:, 0:2], data.iloc[:, 2]
inputs = inputs.fillna(inputs.mean()) # Fill with the mean of the same column
print(inputs)

   NumRooms Alley
0       3.0  Pave
1       2.0   NaN
2       4.0   NaN
3       3.0   NaN
# `pandas` can automatically convert this column into two columns: Alley_Pave and Alley_nan
inputs = pd.get_dummies(inputs, dummy_na=True)
print(inputs)
   NumRooms  Alley_Pave  Alley_nan
0       3.0           1          0
1       2.0           0          1
2       4.0           0          1
3       3.0           0          1

Converting to Tensor Format #

Now that all entries in inputs and outputs are numerical, they can be converted to tensors.

import torch

X = torch.tensor(inputs.to_numpy(dtype=float))
y = torch.tensor(outputs.to_numpy(dtype=float))
X, y

(tensor([[3., 1., 0.],
         [2., 0., 1.],
         [4., 0., 1.],
         [3., 0., 1.]], dtype=torch.float64),
 tensor([127500., 106000., 178100., 140000.], dtype=torch.float64))

Summary #

  1. Write with f.write()
with open(data_file, 'w') as f:
    f.write('NumRooms,Alley,Price\n')  # Column names
  1. Load with pd.read_csv()
import pandas as pd
data = pd.read_csv(data_file)
  1. Positional indexing with .iloc()
inputs, outputs = data.iloc[:, 0:2], data.iloc[:, 2]
  1. Fill with the mean using .fillna()
inputs = inputs.fillna(inputs.mean())
  1. Convert NaN values to indicator columns
pd.get_dummies(inputs, dummy_na=True)
  1. Convert numerical values to tensors with torch.tensor(a.to_numpy(dtype=float))
torch.tensor(outputs.to_numpy(dtype=float))

Related readings


<< prev | Python Data... Continue strolling Python Linear... | next >>

If you want to follow my updates, or have a coffee chat with me, feel free to connect with me: