The code was PHP. All API calls ended like this:
echo json_encode($data) . "\n";
Changed just one character, the period to a comma so the string wasn't duplicated before being output. Problem solved. Felt like a hero.
I thought it was a Big O problem at first, because the code was hacky and used more Arrays that it needed. But it was because it was obtaining all images and then sorting them by time.
I sped it up to milliseconds by making the queries sorted by time.
I.e. changing for (const MyType v : collection) to for (const MyType& v : collection)
Now, of course, there's zero reason to write every debug output with O_SYNC. Classical case of cargo-cult programming. I'd like to say I've never seen something like that again, but then I'd be lying.
The process was loading 40GB files to database and aggregation took more than 5 hours.
Wrote a simple awk one liner with associate array and process was completed in minutes
I took their exact SQL commands and wrapped them in my own scripting, and the simplified single-threaded version was executing in about fifteen minutes.
I went back through and added some explicit parallelization combined with wait commands to ensure that everything in that stage was complete before going to the next stage. That improved version now executes in around 600 seconds.
I changed it to generate single page files instead of one large document (we did 14 days of schedule at a time) and it was finishing in less than 6hr.
Lost that job for that one, the boss had coded the system.
Adding the right indexes in a relational database.
Converting from csv to parquet before querying large datasets on Apache Spark.
tr -cs '[:alnum:]' '[\n*]' | sort | uniq -c
The sort takes a long time (probably just n log n I guess) on a big text. Swapping for
awk '{k[$0]++} END {for (token in k) print token, k[token];}'
and then sorting on the numbers does the same thing faster.
in python. The python script would previously take 4 hours to run. There were lots of small functions and these were called on a loop. I ensured that there were no circular references within any of the functions and then disabled gc (Which is what the gc would look for, variables without circular references would be garbage collected automatically when they go out of scope.)
The script ran in 20 mins.
The hue and cry people raised over the gc.disable though convinced me never to do any unconventional optimizations again.
changed std::map to std::unordered_map