* Iterate with fewer temp allocations * Avoid allocation when not reducing * Avoid slow modulo
There is still more to be done.