NumPy Fundamentals: Why Arrays Beat Python Lists, and How to Use Them

6 min read

pandas, scikit-learn, PyTorch, OpenCV. Peel one layer off any major library in Python’s data ecosystem and you find NumPy arrays. Skip NumPy and learn pandas first, and you will forever handle “why is this operation fast and that loop slow” and “why did the original change too” by feel alone. This post covers NumPy’s five core concepts from first principles. The conclusion up front: NumPy’s essence is one sentence — store same-typed numbers contiguously in memory, and hand the loop to C. Everything else is a corollary.

Why lists are slow: a sea of pointers #

The Python list [1, 2, 3] is not three numbers sitting side by side. It is three pointers to objects, with the actual numbers (int objects) scattered around memory. A loop summing a million list elements pays a million pointer chases, a million type checks, and a million object operations.

ndarray is the opposite. In exchange for accepting the constraint that every element has the same type (dtype), the actual values are packed into contiguous memory. Two things then come free: type checking happens once per array instead of once per element, and the whole loop can run in compiled C. Add CPU cache efficiency, and you get the tens-to-hundreds-of-times gap over lists in numeric work.

array_basics.py
import numpy as np

a = np.array([1.0, 2.5, 3.0])      # dtype inferred as float64
z = np.zeros((3, 4))               # 3 rows, 4 columns of zeros
r = np.arange(0, 10, 2)            # [0 2 4 6 8]
print(a.dtype, z.shape, z.ndim)    # float64 (3, 4) 2

An array’s vitals reduce to three attributes: dtype (element type), shape (size per dimension), ndim (number of dimensions). When NumPy code stops making sense, printing these three is where debugging starts.

Vectorization: not writing loops is the syntax #

How to use NumPy fits in one sentence: do not iterate elements with a for loop; apply operations to whole arrays.

vectorization.py
prices = np.array([12000, 45000, 8900, 23000])

# not this
discounted = [p * 0.9 for p in prices]

# this: vectorization
discounted = prices * 0.9              # multiply everything
over_20k = prices[prices > 20000]      # filter with a boolean mask
total = prices.sum()

prices * 0.9 processes the whole array in one C loop. prices > 20000 produces a True/False array (a boolean mask), and using it as an index extracts only matching elements. Everything you used to do with “loop + if” converts to this pattern. If you find yourself pulling ndarray elements out one at a time inside a Python for loop, you are almost certainly using it wrong — the speed advantage evaporates, and it can end up slower than a list.

Broadcasting: the rule that lets shapes differ #

Multiplying an array by a scalar in prices * 0.9 was the simplest case of broadcasting. The rule: compare the two shapes from the trailing end; each dimension must be equal or one of them must be 1, which gets stretched to match.

broadcasting.py
matrix = np.array([[1, 2, 3],
                   [4, 5, 6]])       # shape (2, 3)
row = np.array([10, 20, 30])         # shape (3,)
col = np.array([[100], [200]])       # shape (2, 1)

matrix + row   # row added to each row  → (2, 3)
matrix + col   # col added to each column → (2, 3)

This rule is why “subtract the column mean from every row” becomes one loop-free line. Incompatible combinations raise errors, so when you hit a shape error, write the two shapes aligned from the right and the cause shows itself.

The trap: slices are views, not copies #

This is where beginners get burned hardest. NumPy slices do not copy data — they create a view onto the same memory.

view_vs_copy.py
a = np.arange(10)     # [0 1 2 ... 9]
b = a[2:5]            # a view: shares a's memory
b[0] = 999
print(a[2])           # 999 — the original changed!

c = a[2:5].copy()     # explicit when you need an independent copy

Python list slices (lst[2:5]) hand you a copy, so list intuition here corrupts your originals. It is a deliberate design for handling huge arrays without copying, so memorize it as a rule: slicing to modify means an explicit .copy(). Note that boolean mask indexing (a[a > 5]) does produce a copy — it is contiguous-range slices that become views.

axis: how to read the direction of aggregation #

From two dimensions up, aggregation has direction — the axis argument — and the rule is single: the axis you specify is the one that disappears.

axis_sum.py
sales = np.array([[10, 20, 30],      # store 1, Jan-Mar
                  [40, 50, 60]])     # store 2, Jan-Mar

sales.sum()          # 210 — everything
sales.sum(axis=0)    # [50 70 90] — fold along rows: monthly totals
sales.sum(axis=1)    # [60 150] — fold along columns: per-store totals

Folding axis=0 of a (2, 3) shape yields (3,); folding axis=1 yields (2,). Thinking “the resulting shape with that axis removed” survives into three dimensions and beyond, where “0 is vertical, 1 is horizontal” collapses. This instinct carries straight into pandas’ axis and the dim of deep learning frameworks.

When NumPy is not the answer #

Know the tool’s boundary too.

  • Small data, tens of elements: array-creation overhead means lists are faster and simpler. NumPy is a tool that wins as data grows.
  • Mixed types, labeled data: if it is not a numeric matrix but “a table with names, dates, amounts,” pandas is right from the start — pandas uses NumPy underneath.
  • String processing, branch-heavy logic: forcing vectorization-hostile logic into NumPy produces unreadable code. Where plain Python is the answer, leave it be.

To confirm performance is actually the problem, measure instead of guessing — profiling tools were covered in py-spy and Advanced Modern Python #7.

Summary #

  • Lists are rows of pointers and slow; ndarray packs one dtype into contiguous memory and gains C loops plus cache efficiency. That is all of NumPy.
  • The core of the syntax is vectorization: whole-array operations and boolean masks instead of for loops. Looping over an ndarray is a warning sign.
  • Broadcasting compares shapes from the trailing end and stretches dimensions of 1. Shape errors resolve by writing the two shapes side by side.
  • Slices are views sharing the original’s memory. Slicing to modify means an explicit .copy().
  • Read aggregation as “the specified axis disappears” — an instinct that carries into pandas and deep learning frameworks.
  • Small data, labeled tables, and branchy logic belong to lists, pandas, and plain Python. NumPy is the tool for large numeric arrays.
X