timerring

Python Linear Algebra

July 14, 2024 · 7 min read
Tutorial
Python
If you have any questions, feel free to comment below. Click the block can copy the code.
And if you think it's helpful to you, just click on the ads which can support this site. Thanks!

Scalar #

Scalar variables are denoted by ordinary lowercase letters (for example, $x$, $y$, and $z$). $\mathbb{R}$ denotes the space of all (continuous) real-valued scalars, and $x\in\mathbb{R}$ is the formal way to say that $x$ is a real-valued scalar. A scalar is represented by a tensor with only one element.

Vector #

A vector can be viewed as a list of scalar values, which are called the vector’s elements or components. Vectors are usually denoted by bold, lowercase symbols (for example, $\mathbf{x}$, $\mathbf{y}$, and $\mathbf{z})$).

$$\mathbf{x} =\begin{bmatrix}x_{1}  \\x_{2}  \\ \vdots  \\x_{n}\end{bmatrix},$$

Length, Dimensionality, and Shape #

$\mathbf{x}\in\mathbb{R}^n$ indicates that the vector $\mathbf{x}$ consists of $n$ real-valued scalars. The length of a vector is often called its dimension.

  • len(x) accesses the length of a tensor.
  • .shape shows the dimensionality of a vector.

When a tensor represents a vector (with only one axis), its length can be accessed through the .shape attribute. Shape is a tuple listing the tensor’s length (dimension) along each axis.

  • The dimension of a vector or axis refers to its length, that is, its number of elements.
  • The dimensionality of a tensor refers to the number of axes it has. In this sense, the dimension of an axis of a tensor is the length of that axis.

Matrix #

Matrices are usually denoted by bold, uppercase letters (for example, $\mathbf{X}$, $\mathbf{Y}$, and $\mathbf{Z}$), and are represented in code as tensors with two axes.

The mathematical notation $\mathbf{A} \in \mathbb{R}^{m \times n}$ denotes a matrix $\mathbf{A}$ consisting of $m$ rows and $n$ columns of real-valued scalars. When a matrix has the same number of rows and columns, it is called a square matrix.

$$\mathbf{A}=\begin{bmatrix} a_{11} & a_{12} & \cdots & a_{1n} \\ a_{21} & a_{22} & \cdots & a_{2n} \\ \vdots & \vdots & \ddots & \vdots \\ a_{m1} & a_{m2} & \cdots & a_{mn} \\ \end{bmatrix}.$$

We can use reshape to form a matrix. A = torch.arange(20).reshape(5, 4)

We can simply use the lowercase letter with subscripts $a_{ij}$ to index matrix $\mathbf{A}$, inserting commas between separate indices only when necessary, such as $a_{2,3j}$ and $[\mathbf{A}]_{2i-1,3}$.

The transpose $\mathbf{a}^\top$ of a matrix swaps its rows and columns.

A.T

Tensor #

A tensor is a general way to describe an $n$-dimensional array with any number of axes. Tensors are denoted by uppercase letters in a special font (for example, $\mathsf{X}$, $\mathsf{Y}$, and $\mathsf{Z}$). Tensors become more important when working with images, which appear as $n$-dimensional arrays.

Basic Properties of Tensor Operations #

Given any two tensors of the same shape, the result of any elementwise binary operation will be a tensor of the same shape

A = torch.arange(20, dtype=torch.float32).reshape(5, 4)
B = A.clone()  # Allocate a copy of A to B by allocating new memory
A, A + B

The elementwise multiplication of two matrices is called the Hadamard product (mathematical symbol $\odot$) $$

\mathbf{A} \odot \mathbf{B} = \begin{bmatrix}     a_{11}  b_{11} & a_{12}  b_{12} & \dots  & a_{1n}  b_{1n} \     a_{21}  b_{21} & a_{22}  b_{22} & \dots  & a_{2n}  b_{2n} \     \vdots & \vdots & \ddots & \vdots \     a_{m1}  b_{m1} & a_{m2}  b_{m2} & \dots  & a_{mn}  b_{mn} \end{bmatrix}.

$$

A * B

Multiplying or adding a scalar to a tensor does not change the tensor’s shape; each element of the tensor is multiplied by or added to the scalar.

Reduction #

  • x.sum() computes the sum of the elements of x. By default, calling the sum function reduces the tensor along all axes, producing a scalar. We can specify which axis of the tensor to reduce by summation: A_sum_axis0 = A.sum(axis=0); similarly, axis=1 sums by column. A.sum(axis=[0, 1]) gives the same result as A.sum ().
  • A.numel() counts the elements.
  • A.mean(), A.sum() / A.numel() computes the mean. The mean function can also reduce a tensor along a specified axis. A.mean(axis=0), A.sum(axis=0) / A.shape[0]

Non-Reduction Sum #

Sometimes it is useful to keep the number of axes unchanged when computing a sum or mean. Since sum_A still has two axes after summing each row, we can (divide A by sum_A through broadcasting).

sum_A = A.sum(axis=1, keepdims=True)
sum_A

tensor([[ 6.],
        [22.],
        [38.],
        [54.],
        [70.]])

sum_A.shape
# torch.Size([5, 1])

A / sum_A

If we want to compute the cumulative sum of the elements of A along an axis, such as axis=0 (across rows), we can call cumsum. This computes cumulative sums.

A.cumsum(axis=0)

tensor([[ 0.,  1.,  2.,  3.],
        [ 4.,  6.,  8., 10.],
        [12., 15., 18., 21.],
        [24., 28., 32., 36.],
        [40., 45., 50., 55.]])

Dot Product #

The dot product is the sum of elementwise products at corresponding positions torch.dot(x, y) == torch.sum(x * y)

Given two vectors $\mathbf{x},\mathbf{y}\in\mathbb{R}^d$, their dot product $\mathbf{x}^\top\mathbf{y}$ (or $\langle\mathbf{x},\mathbf{y}\rangle$) is the sum of elementwise products at corresponding positions: $\mathbf{x}^\top \mathbf{y} = \sum_{i=1}^{d} x_i y_i$.

The dot product is very useful. For example, given a set of values represented by a vector $\mathbf{x} \in \mathbb{R}^d$ and a set of weights represented by $\mathbf{w} \in \mathbb{R}^d$, the weighted sum of the values in $\mathbf{x}$ according to the weights $\mathbf{w}$ can be expressed as the dot product $\mathbf{x}^\top \mathbf{w}$.

When the weights are nonnegative and sum to 1 (that is, $\left(\sum_{i=1}^{d}{w_i}=1\right)$), the dot product represents a weighted average. After normalizing two vectors to unit length, their dot product represents the cosine of the angle between them.

Matrix-Vector Product #

matrix-vector product

Express matrix $\mathbf{A}$ in terms of its row vectors:

$$\mathbf{A}= \begin{bmatrix} \mathbf{a}^\top_{1} \\ \mathbf{a}^\top_{2} \\ \vdots \\ \mathbf{a}^\top_m \\ \end{bmatrix},$$

Each $\mathbf{a}^\top_{i} \in \mathbb{R}^n$ is a row vector representing the $i$th row of the matrix. The matrix-vector product $\mathbf{A}\mathbf{x}$ is a column vector of length $m$ whose $i$th element is the dot product $\mathbf{a}^\top_i \mathbf{x}$. We can view multiplication by a matrix $\mathbf{A} \in \mathbb{R}^{m \times n}$ as a transformation of vectors from $\mathbb{R}^{n}$ to $\mathbb{R}^{m}$. For example, multiplication by a square matrix can represent rotation. [See 3 B 1 B]

We can also use matrix-vector products to describe the complex computations required to solve for each layer of a neural network given the values of the preceding layer. To represent a matrix-vector product with tensors in code, we use the mv function. Calling torch.mv(A, x) for matrix A and vector x performs a matrix-vector product. Note that the number of columns of A (its length along axis 1) must equal the dimension of x (its length).

Matrix-Matrix Multiplication #

Matrix-matrix multiplication can simply be called matrix multiplication and should not be confused with the “Hadamard product”.

For two matrices $\mathbf{A} \in \mathbb{R}^{n \times k}$ and $\mathbf{B} \in \mathbb{R}^{k \times m}$, we can view matrix-matrix multiplication $\mathbf{AB}$ as performing $m$ matrix-vector products and concatenating the results into an $n \times m$ matrix. In short, rows on the left and columns on the right.

B = torch.ones(4, 3)
torch.mm(A, B)

Norms #

Some of the most useful operators in linear algebra are norms. Informally, the norm of a vector indicates how large that vector is. (The concept of size does not concern dimensionality, but rather the magnitude of its components.)

In linear algebra, a vector norm is a function $f$ that maps a vector to a scalar. Given any vector $\mathbf{x}$, a vector norm must satisfy several properties.

  1. If we scale all elements of a vector by a constant factor $\alpha$, its norm scales by the absolute value of the same factor: $$f(\alpha \mathbf{x}) = |\alpha| f(\mathbf{x}).$$
  2. The triangle inequality: $$f(\mathbf{x} + \mathbf{y}) \leq f(\mathbf{x}) + f(\mathbf{y}).$$
  3. A norm must be nonnegative: $$f(\mathbf{x}) \geq 0.$$ In most cases, the smallest possible size of anything is 0.
  4. A norm attains its minimum of 0 if and only if every component of the vector is 0. $$\forall i, [\mathbf{x}]_i = 0 \Leftrightarrow f(\mathbf{x})=0.$$

Norms sound much like measures of distance. Euclidean distance and the notions of nonnegativity and the triangle inequality in the Pythagorean theorem may offer some intuition.

In fact, Euclidean distance is an $L_2$ norm: suppose the elements of an $n$-dimensional vector $\mathbf{x}$ are $x_1,\ldots,x_n$. Its $L_2$ norm is the square root of the sum of the squares of its elements:

$$\|\mathbf{x}\|_2 = \sqrt{\sum_{i=1}^n x_i^2},$$

Here, the subscript $2$ is often omitted from the $L_2$ norm, so $|\mathbf{x}|$ is equivalent to $|\mathbf{x}|_2$. In code, we can compute the $L_2$ norm of a vector as follows.

u = torch.tensor([3.0, -4.0])
torch.norm(u)
# tensor(5.)

There is also the $L_1$ norm, which is the sum of the absolute values of the vector’s elements:

$$\|\mathbf{x}\|_1 = \sum_{i=1}^n \left|x_i \right|.$$

Compared with the $L_2$ norm, the $L_1$ norm is less affected by outliers. To compute the $L_1$ norm, we combine the absolute value function with elementwise summation.

torch.abs(u).sum()

The $L_2$ norm and $L_1$ norm are both special cases of the more general $L_p$ norm:

$$\|\mathbf{x}\|_p = \left(\sum_{i=1}^n \left|x_i \right|^p \right)^{1/p}.$$

Like the $L_2$ norm of a vector, a matrix also has a norm. The ** Frobenius norm of $\mathbf{X} \in \mathbb{R}^{m \times n}$ is the square root of the sum of the squares of its elements:**

$$\|\mathbf{X}\|_F = \sqrt{\sum_{i=1}^m \sum_{j=1}^n x_{ij}^2}.$$

The Frobenius norm satisfies all the properties of a vector norm; it is like the $L_2$ norm of a matrix-shaped vector. Calling the following function computes the Frobenius norm of a matrix.

torch.norm(torch.ones((4, 9)))

Related readings


<< prev | Python Data... Continue strolling Housewarming... | next >>

If you want to follow my updates, or have a coffee chat with me, feel free to connect with me: