Preface

Today I watched 3b1b1’s ‘Essence of Linear Algebra]2’, which gave me many new insights and clarified many things I hadn’t fully understood before (or rather, formulas I had blindly memorized just to pass exams). I couldn’t wait to write this article. The geometric intuitions used here do not involve formal mathematical proofs.

Knowledge

Matrix Multiplication

First, let me record some key points mentioned in the video. There's nothing much to say about vectors themselves; instead, vector transformations lead us to matrix multiplication. Vector transformations are essentially coordinate system transformations, or what we call linear transforms. Consider a vector $\overrightarrow{v}=[x\hat{i}, y\hat{j}]^{T}$. Applying a transformation $A=\begin{bmatrix}a & c\\ b & d\end{bmatrix}$ to it: $\hat{i},\hat{j}$ represents the basis vectors. The transformation applied to $\hat{i}$ is $\begin{bmatrix}a\\b\end{bmatrix}$, and the transformation applied to $\hat{j}$ is $\begin{bmatrix}c\\d\end{bmatrix}$. The coordinates after the transformation are $\begin{bmatrix}a & c\\ b & d\end{bmatrix}\begin{bmatrix}x\\ y\end{bmatrix}=x\begin{bmatrix}a\\ b\end{bmatrix}+y\begin{bmatrix}c\\ d\end{bmatrix}=\begin{bmatrix}ax+cy\\bx+dy\end{bmatrix}$

Finally, I don’t have to memorize the row-column order for matrix multiplication!

This extension to higher dimensions works the same way. Below, let’s consider multiple transformations. Applying two transformations $A,B$ (which is equivalent to $AB\overrightarrow{v}$) is essentially a composite function $A(B\overrightarrow{v})$: first performing a $B$ transformation, followed by a $A$ transformation. Why can’t we swap this order? A more intuitive explanation is that after the first $B$ transformation, the coordinate system has already changed. Although $A$ transforms coordinates in the same way, its effect on the already-transformed $B\overrightarrow{v}$ is different from its effect if applied directly to $\overrightarrow{v}$. This explains why matrix multiplication is not commutative.

What about associativity? Notice that both $(AB)C\overrightarrow{v}$ and $A(BC)\overrightarrow{v}$ are actually associated from right to left. The result of multiplying the first two matrices is also combined in order, so the associative law does hold (even if the intuitive feeling might seem strange here, this is indeed the correct transformation sequence).

It is my experience that proofs involving matrices can be shortened by 50% if one throws the matrices out. - Emil Artin

Determinant

Previously, my learning of determinants was biased; I only used them to solve systems of linear equations and never considered their geometric meaning. Here, the geometric meaning of a determinant is the factor by which the unit area changes under a ’transformation’—essentially, the ‘area’ of the unit square after the transformation.

This actually made me pause for a moment. If we accept the above statement, then the area after the matrix transformation represented by $[1,1]^{T}$ should be the determinant. Following this logic, $\begin{bmatrix}a & c\\b & d\end{bmatrix}\begin{bmatrix}1\\1\end{bmatrix}=\begin{bmatrix}a+c\\b+d\end{bmatrix}$, the area of the identity matrix should be $(a+c)(b+d)=ab+ad+bc+cd$. However, we all know the value of a 2x2 determinant should be $ad-bc$. Why the discrepancy? It's simply because 'the coordinate system changed.' $ad-bc$ uses the original coordinate system, while $(a+c)(b+d)$ uses the transformed one.

So, can a determinant be negative? Yes, this corresponds to a flip in the coordinate system. For example, rotating the plane coordinate system $x$ to $y$ by 90° counterclockwise, or changing the $xyz$ coordinate system from a right-handed system to a left-handed system, will both result in a negative determinant. For higher dimensions, fixing all other dimensions and transforming two of them will make the determinant negative; transforming another dimension makes it positive again, then negative again, and so on. However, this seems impossible to prove geometrically, as the human brain struggles to visualize high-dimensional spaces.

What if the determinant of a transformation becomes 0? The unit area becomes 0, meaning the transformation has been ‘dimensionally reduced’ into a line. This implies that the column vectors of the transformation are not linearly independent.

At the same time, 3b1b left a thought-provoking question $\det(M_1M_2)?=\det(M_1)\det(M_2)$. The result is, of course, equal. 3b1b did not provide a geometric explanation, but here is my understanding: it is essentially multiplication. Looking at the right side, following the right-to-left association order, we first scale the original unit matrix by $\det(M_2)$, then apply the $M_1$ transformation to scale it to $\det(M_1)$, obtaining the final unit area. This is equivalent to performing two transformations sequentially, reorganizing the coordinates after each transformation before proceeding to the next. On the left side of the equation, it’s as if both transformations were performed at once, directly scaling to $\det(M_1)\det(M_2)$, thus yielding $\det(M_1M_2)=\det(M_1)\det(M_2)$.

Inverse Matrix & Rank

The inverse matrix, or inverse “transformation,” brings us to a point where all matrices can essentially be referred to as “transformations,” because what we are dealing with here are coordinate system transformations, i.e., linear transforms. One point I previously forgot to mention is that the result of a coordinate system transformation caused by a linear transform is always linear, meaning the coordinate axes remain “straight,” must be “parallel,” and are also “equidistant.” If these conditions are not met, a straight line in the original coordinate system would no longer be a straight line, and thus it would no longer be a linear transformation.

As an inverse transformation, it is written as $A^{-1}$. Applying a transformation and then transforming it back is essentially $A^{-1}A$ (which might explain why we always write $A^{-1}$ on the left? Because the combination always starts from the right). Solving an equation $Ax=v$ is essentially finding a transformation that maps $v$ to $x$. Note that it is not transforming $x$ into $v$, because $x$ is the unknown variable. This is essentially $A^{-1}Ax=A^{-1}v$, and the solution is $x=A^{-1}v$.

If $\det(A)=0$, then the transformation that has been subjected to “dimensional reduction” can never be restored to its pre-transformation state. This is equivalent to saying you can never find a plane from two parallel line vectors, meaning no inverse exists, yet there are infinitely many solutions.

After transformation, the minimum dimension it compresses to is the rank. This is derived from the transformation of each column vector, resulting in the column space.

Regardless of any transformation, the vector $\begin{bmatrix}0\\0\end{bmatrix}$ must remain at the origin. For a full-rank matrix, only the $[0,0]^{T}$ vector stays at the origin. For a non-full-rank matrix, a series of vectors will be compressed until they reach the $[0,0]^{T}$ vector; these constitute the "null space" or "kernel."

Non-square Matrices

Considering the “column space,” left-multiplying by a non-square matrix is essentially performing a projection. If you left-multiply by a square matrix of size $3\times 2$, it is essentially transforming and projecting a 2D plane into 3D space.

The dot product itself has geometric meaning; it can also be viewed as the projection described above, projecting onto a one-dimensional space vector.

More thing

Let’s think about other expressions within the matrix.

  1. $(AB)^{-1}=B^{-1}A^{-1}$ .

    From right to left, first combine the $B$ transformation, then combine the $A$ transformation. To reverse this process, considering that the coordinate system must remain stationary, we must first apply the inverse of the $A$ transformation, and then apply the inverse of the $B$ transformation.