8,384 karma · joined December 3, 2010
I like automating things. Currently building lots of software in Python / Typescript / Postgres.
Contact me: aidankane on that google hosted email platform.
If you were reading it uncharitably, you might say that it seems like the authors don't understand Postgres that well. A more charitable reading is it's presented as a journey of discovery that they expect their readers to have been on, but it's hard to tell.
> Well, now when we have this job done and finally open the log file to find out… a long wall of plain text. It actually gives us no info about the query plan except for its execution time.
As someone who has spent a fair amount of time with the PG planner there were a bunch of things in there that felt off. Eg in the below my guess would be that this is to do with paging, not statistics.
> By the time you retrieve the query from log files and run it again, the dataset and statistics will be completely different. So when you finally execute the query for debugging, it might run quickly this time, which makes you wonder why it was slow before.
If someone is new to query optimisation in PG this article isn't a great place to start. There's a lot of content that reads like the author doesn't know Postgres well but, given that they've written a tool to visualise pg queries, I'm going to assume that they do and instead the article just doesn't quite hit the mark.
If you’re working with Postgres you owe it to yourself to understand how to read query planner output.
I imagine those out there in the field have to think about this a lot more than the rest of us.
Edit: removed link as was tangentially related
I played around with warm sqlite too. That was really nice but I decided against it due to the fact that it was totally unsupported.
I found that it was pretty easy to get cursor into a place where it had decimated the file and each tab complete suggestion became more and more broken.
Claude was given the task wholesale and did a reasonable job but introduced a subtle bug by moving a tracking call outside of the Ajax promise and I could not convince it to put it right. It kept apologising and then offering up more incorrect code.
I’d say that the original result was good enough that I could pretty much take it and fix it, but only because I knew all the code and libraries well enough. It was only about 150 lines of simple code and by the time I’d finished I was joking with the team that I could have spent all the time wrestling vim macros instead and come out about the same.
What’s your experience been with correctness?
Different people have different requirements.
On my local machine it’s even better because I can run the code, let it break and then jump straight into the debugger where I can move up and down through the stack. I can sit in the context at any point and play with the data. I can call any Python function I want to see how the code behaves.
I can see the advantage in terms of just needing a single tuple for a reads. So a timestamp + value model would likely take twice as much heap space than your approach?
Given that you’re probably always just inserting new data you could use a brin index to get fast reads on the date ranges. Would be interesting to see it in action and play around to see the tradeoffs. The model you’ve settled on sounds like it would be a pain to query.
https://github.com/aidos/dotfiles/blob/master/_config/ipytho...
I live in my Python REPL (ipython) and I couldn’t live without my hot reloading without losing state. Allows for hacking about a bit directly in the REPL. Once something starts to take shape you can push it into a file, import it into your REPL then hack about with it in a text editor.
The hot reloading is great when you have a load of state in classes and you can change code in the editor and have that immediately updated on your existing objects in the REPL.
Have seen quite a few clips of SailGP but am yet to watch a full race. Will give it a bit more of a look.
Here’s a clip that gives a bit of an idea.
These boats are amazing and I’ve been following along this year but it’s the first time I’ve found it harder to connect. Now the team and head down cycling it’s hard to feel the same way as when they were running about the boat.
I’ve sailed a fair amount and love the sport. I love all the boats, monohull, cat, flying, in the water. If you haven’t seen it, check it out. The technical mastery that has created these machines is off the chart. The fact that with just the right shape, you can harness the breeze like they do just defines belief.
https://gist.github.com/aidos/5a6a3fa887f41f156b282d72e1b79f...
For anyone else, here's the awk for combining lines in the log files for making them greppable too: https://gist.github.com/aidos/44a9dfce3c16626e9e7834a83aed91...
I dump about 150GB of Postgres logs a day (I know, it's over the top but I only keep a few days worth and there have been several occasions where I was saved by being able to pick through them).
At that size you even need to give up on grepping, really. I've written a tiny bash script that uses the fact that log lines start with a timestamp and `dd` for immediate extraction. This allows me to quickly binary search for the location I'm interested in.
Then I can `dd` to dump the region of the file I want. After that I have an little awk script that lets me collapse the sql lines (since they break across multiple lines) to make grepping really easy.
All in all it's a handful of old school script that makes an almost impossible task easy.
That’s the problem tailwind solves.
Coldfusion had a bad rep but it was actually a pretty good tool to work with. Railo made it a million times better in every way - the performance was solid and the environment was full featured enough to let you focus on the code.
The original adobe (macromedia, right?) version was a mess. There were no functions but you could import modules. Massive caveat that those modules were allowed to mutate state outside of themselves so people would import them to, for example, run some db queries and set some variables in the calling code with the results. You can imagine the absolute unmaintainable mess that ensued.
I revert to “group by 1, 2, 3… “ when I’m just hacking about. Group by all would definitely be an improvement.
Off the back of this comment I’ve just flicked back to a random video of my then 2 yo where we have a discussion about how houses in real life aren’t normally the colour of the ones in the kids book she’s reading.
mutool clean -d in.pdf out.pdf
At that point you’ll realise that a PDF is mostly just a list of objects and that those objects can reference each other. After that you’ll journey through the spec understanding what each type of object does and what the fields in it control. The graphics stream itself is just a stack based co-ordinates drawing system that’s easy to follow too.By way of an example. Here's an object that represents a Page. You can see the dimensions in the MediaBox. The contents themselves are contained at object "9 0 obj" ("9 0 R" is the pointer to it):
2 0 obj
<<
/Type /Page
/MediaBox [ 0 0 612 792 ]
/Contents 9 0 R
>>
endobj
Meanwhile "9 0 obj" has the drawing instructions. They seem a little weird at first glance but you see the values ".23999999 0 0 -.23999999 0 792" each get pushed on the stack and then "cm" pops them to interpret them as the transformation matrix. 9 0 obj
<<
/Length 18266
>>
stream
.23999999 0 0 -.23999999 0 792 cm
q
0 0 2551 3301 re
...
The depth and detail of all of the different possible things that can be represented in a PDF is insane. But understanding the structure above is all you need to begin your journey!EDIT The rest of your journey is contained in this epic document: https://opensource.adobe.com/dc-acrobat-sdk-docs/pdfstandard...