HNHacker News
TopNewBestAskShowJobs

clappski

443 karma · joined February 29, 2016

C++ software developer, based in London.

me <at> jakeclapp <dot> com

submissionscomments
clappski··on HTML over WebSockets: real-time SPAs with barely any JavaScript
I’ve had more difficulty building http1.1 implementations that broadly work, still a lot of weird stuff like incorrect chucked message parsing you have to handle on the server side. Yeah plain TCP is obviously simpler but you can’t consume a raw socket in a browser, that’s why web sockets win for me - easy to consume over api and browser.
clappski··on HTML over WebSockets: real-time SPAs with barely any JavaScript
You can genuinely implement the protocol with good performance in 150/200 lines, from memory;

1. Read the opcode

2. If PING, send PONG. If binary or text, continue

3. Read the header size

4. Read the payload size (which is a variable-width integer, hence getting the header size first)

5. Unmask the payload

6. If fin in the header is false, add the payload to a buffer. Else, either immediately dispatch (continuation buffer is empty) or buffer and dispatch (continuation buffer isn’t empty)

There are extensions like setting RSV1 to indicate you need to decompress the payload, but that’s the only one I’ve encountered building clients for >20 WebSocket based APIs.

clappski··on HTML over WebSockets: real-time SPAs with barely any JavaScript
In my experience (building things like realtime trade monitoring front ends), client side processing sucks. You have 10 or 20 data sources that need joining up in a client defined set of ways (think viewing X price feed with Y trade feed and some Z meta data feed), all natively sending many updates a second - I’ve tried and failed to build a client side heavy implementation (although that might just be my lack of front end skills!), much better to have some DAG of backend processing units that the front end can consume feeds from and remain really light, just managing subscriptions and rendering the data.
clappski··on HTML over WebSockets: real-time SPAs with barely any JavaScript
The claim I was rebutting was specifically about latency. Of course you can make a mess of it, no different than you can with request/response based protocols.

On the other hand, you are going to struggle to implement bidirectional event based protocols without WebSockets that can be consumed by browser based clients or an application - WebSockets are very useful for that exact case, where you have a mixed client type consuming. It means you don’t need to offer web hooks and rest and some streaming protocol, you just have a WebSocket with some defined JSON message types without the rigmarole of distributing .proto or whatever.

clappski··on HTML over WebSockets: real-time SPAs with barely any JavaScript
Yes, WebSockets are very simple to implement a server or a client for. It’s a small variable size header, 5 or so frame types and a typically constant mask over the contents.
clappski··on HTML over WebSockets: real-time SPAs with barely any JavaScript
> The latency is the same because modern browsers multiplex HTTP requests over a single TCP connection that is left open.

In my experience this isn’t true; firstly you’re relying on an implementation detail of the platform that you’re executing on, of which you have no control over on the client side. Secondly, even if you aren’t opening a new connection per request, you’re still travelling through an entire HTTP stack implementation rather than the incredibly simple WebSocket protocol - effectively a length and a mask to get the contents, rather than some (in http1 land) fuzzy parser.

If you can guarantee you’re hitting http2 or http3 then you might be closer in latency, but due to the complexity of both I would imagine plain http1 negotiated persistent WebSockets provide the best latency.

clappski··on Learn SQL Once, Use It for 30 Years
At least in that case you can refactor the stored proc to be more performant without pushing application changes.
clappski··on David Attenborough's 100th Birthday
Not at all, for example at secondary/high school you would address male teachers interchangeable as ‘sir’ and ‘mr. Last name’.
clappski··on I dumped Windows 11 for Linux, and you should too
The mistake is that you did a system update when you wanted to use the computer. Not implying that system updates should be a dangerous thing to do, but just something learnt from experience - I’ve had similar issues, especially with Nvidea drivers and kernel versions getting updated at the same time. The take away is keep the updates to when you have an hour to debug or get comfortable rolling back updates.
clappski··on Golang's big miss on memory arenas
I like the priorities.

I think a core thing that's missing is that code that performs well is (IME) also the simplest version of the thing. By that, I mean you'll be;

- Avoiding virtual/dynamic dispatch

- Moving what you can up to compile time

- Setting limits on sizing (e.g. if you know that you only need to handle N requests, you can allocate the right size at start up rather than dynamically sizing)

Realistically for a GC language these points are irrelevant w.r.t. performance, but by following them you'll still end up with a simpler application than one that has no constraints and hides everything behind a runtime-resolved interface.

clappski··on Stdio(3) change: FILE is now opaque
You'd have an error about an incomplete type - see https://godbolt.org/z/G4Gsfn7MT

> a pointer, whose size is always known

Yeah, this is exactly how it works. You work with a pointer that acts like a void* in your code, and the library with the definition is allowed to reach into the fields of that pointer. Normally you'd have a C API like

    struct Op;
    Op* init_op();
    void free_op( Op* );
    void do_something_with_op( Op* );

in the header provided by the library that you compile as part of your code, and the definition/implementation in some .a or .so/.dll that you'll link against.*
clappski··on Stdio(3) change: FILE is now opaque
When we're talking about opaque it's really in relation to an individual translation unit - somewhere in the binary or its linked libraries the definition has to exist for the code that uses the opaque type.
clappski··on Why does unsafe multithreaded std:unordered_map crash more than std:map?
Reminds me of the shadow paging used in LMDB, which is effectively the same but at the page level rather than the whole tree and allows each reader their own context to read from, rather than a shared reader context.
clappski··on Python 3.12.0rc1

     f" something { my_dict['key'] } something else "
This works in Python already, allowing for nesting is a big QoL improvement
clappski··on A practical comparison of build and test speed between C++ and Rust
C++ has two types of polymorphism;

- Templates (compile time), which are generic bits of code that are monomorphized over every combination of template parameters that they're used with.

- Virtualization (runtime), classic OO style polymorphism with a VTable (I think this similar to Rust's dyn trait?).

clappski··on Tesla shares tank after U.S. discounts doubled on key models
> haven't seen depth info past the best bid/offer (maybe exists? I'm not an expert)

Most exchanges will provide arbitrary depth in way of a pure order feed. From that you can construct your price book, which gives you the levels of depth.

Some just expose the price book, some price book and order feed, some just order feed.

Some anonymize the order feed, some don't (I think typically in equities it's anonymized although not my asset class, but other markets e.g. power need to be de-anonymized and you can see the other buyers and sellers submissions).

However, that data doesn't represent the market price - I would look to constructing the fair price using the trades made on the exchange rather than the outstanding orders. From that you can create different lenses to view the fair price - e.g. volume weighted, time weighted, other types of averages.

clappski··on The sinister attempts to ‘decolonise’ mathematics
So are you arguing that genocide and slavery aren't objectively bad?
clappski··on The Problem with Go
By sharing the load, we have a #reviews channel that anyone can put something in to get peer review. If there's something that needs attention from a specific person then arranging a call to walk through is a good approach where both reviewee and reviewer can negotiate a time to review.
clappski··on The pool of talented C++ developers is running dry
There's much more reason to do a greenfield project in C++ than Rust - experienced C++ hiring is still considerably easier! Not everything has a purely technical motivator.
clappski··on The forty-year programmer
> Is it boring? Was I right about that?

> who uses programming as a problem solving tool

Programming can be boring, software development is much more exciting!

All software is solving a problem, same as the software that’s the output of the programming you do.

Those problems might be purely technical (like virtualising different CPU architectures) or focussed on improving the efficiency that others can solve problems (like the software running Stack Exchange) or something completely different.

Solving all of those problems requires more than programming, same as the problems you’re trying to solve need more than programming to fully resolve. Building for maintainability, reliability etc. requires much more than mindlessly programming the software.

But even the programming itself is interesting, especially when solving difficult problems.

clappski··on Developers who quit the industry. Why? And what do you do now?
The ship had sailed on

> WFH is a rare, once a month kind of thing

years ago for tech in London, way before 2020 for most businesses.

You'd be better off looking outside of London in the commuter belt or further out if you actually wanted to be in an office everyday, where everyone else is also in an office everyday. Or work in a highly regulated industry - defense, something governmental. Or something that requires you to develop against physical hardware.

clappski··on Building a Cloud Database from Scratch: Why We Moved from C++ to Rust
In practise shared_ptr should only be used when you need to actually share ownership of some dynamically allocated object. You get around it by using it in the right place. You’re correct on the refcounting being a performance issue.

Codebases that use shared_ptr for everything because they think it ‘solves’ having to manage memory are smelly, it shows that no one has thought about the ownership model. Not to say there isn’t a solid use case for it, but some developers use it by default because they don’t fully understand the trade offs/why it’s important to think about ownership of managed memory.

As an example in our medium to large code base we have two places we use a shared_ptr, completely off the hot path and for long lived objects that aren’t copied more than a few times, but need to share ownership of the pointee - in a similar use case of your use of Arc.

Almost all of the time unique_ptr, with a single owner, is exactly what you want to replace the memory management side of a dynamically allocated object. From there you will pass around raw pointers or references to the object managed by the unique_ptr (for functions that need to ‘borrow’ the object).

Something interesting about shared_ptr is how it relates to weak_ptr, although I’ve never seen weak_ptr in anything out side documentation!

clappski··on Python Standard Library changes in recent years
Sometimes you might use bits in an integer to encode data, e.g.;

- A bitset index

- Bitflags (e.g. to represent logging levels)

- Feature flags

`bit_count` (aka `popcount`) gives you an efficient way to figure out how many bits within those structures are set and not set.

clappski··on NY energy grid: Real-time dashboard
Something to bare in mind when you're looking at both of those is the price is highly dependent on a number of factors;

- Type of generation, looking at the NY data they have a huge hydro generation, obviously the UK grid has much less of that

- Time of day, the Elexon graph is showing you the system price per 30 minute settlement period. You can see that some of those periods weren't anywhere near what you're quoting (there was literally a period where the system price was 0£ in the last 24h). https://www.bmreports.com/bmrs/?q=balancing/systemsellbuypri... makes the volatility much more obvious.

- The weather, it plays a big part in the system price due to price disparity of renewables (commonly wind in the UK grid) compared to oil based and gas generation.

- The interconnect, the UK grid has interconnectors to the EU grid so there's some price impact from that

clappski··on Python built-ins worth learning (2019)
Wouldn't `Count` ( https://docs.microsoft.com/en-us/dotnet/api/system.collectio... ) be the preferred way to check if a collection has values? `someCollection.Any()` is the same as Python's `any`
clappski··on Const all the things?
> But wouldn't the ability to restrict arbitrary changes to a struct still be important?

Typically immutability in C++ is at the interface scope rather than structural, e.g. prefer this

    struct MyType {
        int x_;

        MyType( int x ) : x_( x ) {}
    };
   
    void do_the_thing( const MyType& mt ); // mt.x_ = 1; doesn't compile, mt is const
over this

    struct MyType {
        const int x_;

        MyType( int x ) : x_( x ) {}
    };
   
    void do_the_thing( MyType mt ); // mt.x_ = 1; doesn't compile, x_ is const
clappski··on Do Not Log
My comment wasn’t about buffered writing in a main thread, it was about moving the work out of your threads actually executing your program and having them push raw log actions into a queue for another thread to do the formatting and writes to file from.

If you’re logging so much that you’re saturating a dedicated thread that just reads from a queue in shared memory, does log formatting and writes to a file then you have bigger problems than figuring out how to log because there isn’t going to be a solution that lets you log with a deterministic latency without skipping events or wrapping around your queue.

Of course, the approach scales fine if you can have multiple log files per process - e.g. a normal log file and a message log or transaction log, because you can give each file its own queue and thread.

clappski··on Do Not Log
> logging to a synchronized output

There's your performance issue, logging doesn't need to be persisted or even processed by your main program thread or any thread/process actually doing the work. Who's logging straight to stdout or directly ::write'ing a file in a production application?

clappski··on Almost Always Unsigned
What you should be doing for checked arithmetic with GCC is use the builtins for those that aren’t aware;

https://gcc.gnu.org/onlinedocs/gcc/Integer-Overflow-Builtins...

clappski··on Constructors and evil initializers in C++
It’s only possible to do that if your positive integers are passed via template parameters.

All your type traits/SFAINE/concepts/static_asserts happen at compile time, you can’t use them to do any checks on the value of a type, unless it’s a value available at compile time.

Page 1 of 7Next →