NumPy library
NumPy is an open-source project, with fiscal sponsorship from the nonprofit NumFOCUS.
NumPy (imported as np) is a widely used library for fast numeric arrays — the foundation nearly every other data or scientific library in Python is built on. It's a third-party package, not part of the standard library. A NumPy ndarray looks similar to a list, but every element is the same type and math operations apply to the whole array at once, instead of one item at a time.
Setup
pip install numpy
np is the near-universal alias for NumPy — used throughout this page and in virtually every codebase that imports it.
import numpy as np
| Type | Holds | Math operations |
|---|---|---|
list |
Any mix of types | Element-by-element, usually with a loop |
ndarray |
One type, fixed size | Applied to the whole array at once ("vectorized") |
Creating arrays
np.array() builds an ndarray from an existing list — every value gets converted to the same type.
import numpy as np
lengths_ft = np.array([4.5, 12, 8, 6])
print(lengths_ft)
print(lengths_ft.dtype)
Building arrays without a list
np.zeros(n) builds an array of n zeros as a starting point to fill in later. np.arange(stop) counts up from 0 to (but not including) stop, just like the built-in range() — with an optional start and step, exactly like range() too.
np.zeros(4) # array([0., 0., 0., 0.])
np.arange(4) # array([0, 1, 2, 3])
np.arange(0, 10, 2) # array([0, 2, 4, 6, 8])
Run a creating arrays example
All the examples above, combined into one script:
import numpy as np
lengths_ft = np.array([4.5, 12, 8, 6])
print(lengths_ft)
print(lengths_ft.dtype)
import numpy as np
print(np.zeros(4))
print(np.arange(4))
print(np.arange(0, 10, 2))
Array operations
A math operation on an array applies to every element at once — no loop required, and considerably faster than looping over a plain list.
import numpy as np
lengths_ft = np.array([4.5, 12, 8, 6])
lengths_m = lengths_ft * 0.3048
print(lengths_m)
For efficiency, use vectorized operations instead of a Python loop
| Time | Space | |
|---|---|---|
Python for loop |
O(n) | — |
| Vectorized | O(n) (same class, smaller constant) | — |
A Python for loop over a list and a vectorized NumPy operation both touch every element once — O(n) either way, the same Big O class.
The speed difference is a constant factor, not the order of growth: each pass of a Python loop pays the interpreter's per-iteration overhead, while a vectorized operation runs its loop once, in compiled C, underneath a single Python call. That overhead is small per element but adds up — the larger the array, the bigger the gap, even though neither approach's growth rate has changed.
See Efficiency for why this distinction matters.
Aggregating an array
Collapses an entire array down to a single summary number. .mean(), .max(), .min(), and .sum() — the same idea as Python's built-in sum() and max(), but computed directly on the array without converting it back to a list first.
lengths_ft = np.array([4.5, 12, 8, 6])
lengths_ft.mean() # 7.625
lengths_ft.max() # 12.0
lengths_ft.sum() # 30.5
Filtering with a boolean mask
Comparing an array to a number produces a same-size array of True/False values — a boolean mask. Indexing the array with that mask keeps only the elements where it's True. This is the standard way to filter a NumPy array, instead of writing an explicit loop with an if inside it.
lengths_ft = np.array([4.5, 12, 8, 6])
lengths_ft > 7 # array([False, True, True, False])
lengths_ft[lengths_ft > 7] # array([12., 8.])
Run an array operations example
All the examples above, combined into one script:
import numpy as np
lengths_ft = np.array([4.5, 12, 8, 6])
lengths_m = lengths_ft * 0.3048
print(lengths_m)
import numpy as np
lengths_ft = np.array([4.5, 12, 8, 6])
print(lengths_ft.mean())
print(lengths_ft.max())
print(lengths_ft.sum())
import numpy as np
lengths_ft = np.array([4.5, 12, 8, 6])
mask = lengths_ft > 7
print(mask)
print(lengths_ft[mask])