SyncAI.news, a Varaisys broadcasting
3 Numba Tricks for Python Runtime Optimization
MM

Matthew Mayo

· 1 min read

EngineeringKDnuggets

3 Numba Tricks for Python Runtime Optimization

Numba compiles a numeric Python loop to machine code without ever having to leave your Python environment, rewrite anything in C, or vectorize some chunk of code that does not want to be vectorized. When Numba code disappoints, it's nearly never the compiler. The offender usually ends up being the boundary around the compiled code: not crossing it, not making it wide enough, or crossing it during every run. Here are three tricks, all of them the same question asked three ways.

Note that everything below was checked against Numba 0.67.0.

pip install numba

Trick 1: Compiling the Loop Instead of Interpreting It

The baseline is a reduction over a NumPy array, and this is slow for the ordinary reason: the interpreter dispatches on types once per element, ten million times. The decorated version of the code differs by exactly one line. Numba reads the types on the first call to total_jit(), compiles a specialization for them, and every call after that it just runs native code:

Output:

Numba first call (includes compilation): 0.2799 s

Method                 Best (s)   Mean (s)    Speedup  Result
------------------------------------------------------------------------
Plain Python loop        3.0789     3.0789       1.0x  3641603.675817
Numba @njit              0.0384     0.0385      80.2x  3641603.675817
NumPy vectorized         0.0552     0.0606      55.7x  3641603.675816

Results match: True

What hasn't changed is the constraint underneath it. Nopython mode (@njit) "produces much faster code, but has limitations," and those limitations are important. Using Numba well is mostly a matter of keeping the hot function inside the subset of Python and NumPy it can assign types to.

Trick 2: Spreading the Loop Across Every Core

from numba import njit, prange

@njit(parallel=True)
def total_parallel(x):
    total = 0.0
    for i in prange(x.shape[0]):
        total += np.sqrt(x[i]) * np.sin(x[i])
    return total

And modify our results in order to run the new experiment as follows:

Original source

This story was published by KDnuggets and written by Matthew Mayo. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on kdnuggets.com

Similar News