443 karma · joined February 29, 2016
me <at> jakeclapp <dot> com
1. Read the opcode
2. If PING, send PONG. If binary or text, continue
3. Read the header size
4. Read the payload size (which is a variable-width integer, hence getting the header size first)
5. Unmask the payload
6. If fin in the header is false, add the payload to a buffer. Else, either immediately dispatch (continuation buffer is empty) or buffer and dispatch (continuation buffer isn’t empty)
There are extensions like setting RSV1 to indicate you need to decompress the payload, but that’s the only one I’ve encountered building clients for >20 WebSocket based APIs.
On the other hand, you are going to struggle to implement bidirectional event based protocols without WebSockets that can be consumed by browser based clients or an application - WebSockets are very useful for that exact case, where you have a mixed client type consuming. It means you don’t need to offer web hooks and rest and some streaming protocol, you just have a WebSocket with some defined JSON message types without the rigmarole of distributing .proto or whatever.
In my experience this isn’t true; firstly you’re relying on an implementation detail of the platform that you’re executing on, of which you have no control over on the client side. Secondly, even if you aren’t opening a new connection per request, you’re still travelling through an entire HTTP stack implementation rather than the incredibly simple WebSocket protocol - effectively a length and a mask to get the contents, rather than some (in http1 land) fuzzy parser.
If you can guarantee you’re hitting http2 or http3 then you might be closer in latency, but due to the complexity of both I would imagine plain http1 negotiated persistent WebSockets provide the best latency.
I think a core thing that's missing is that code that performs well is (IME) also the simplest version of the thing. By that, I mean you'll be;
- Avoiding virtual/dynamic dispatch
- Moving what you can up to compile time
- Setting limits on sizing (e.g. if you know that you only need to handle N requests, you can allocate the right size at start up rather than dynamically sizing)
Realistically for a GC language these points are irrelevant w.r.t. performance, but by following them you'll still end up with a simpler application than one that has no constraints and hides everything behind a runtime-resolved interface.
> a pointer, whose size is always known
Yeah, this is exactly how it works. You work with a pointer that acts like a void* in your code, and the library with the definition is allowed to reach into the fields of that pointer. Normally you'd have a C API like
struct Op;
Op* init_op();
void free_op( Op* );
void do_something_with_op( Op* );
in the header provided by the library that you compile as part of your code, and the definition/implementation in some .a or .so/.dll that you'll link against.* f" something { my_dict['key'] } something else "
This works in Python already, allowing for nesting is a big QoL improvement- Templates (compile time), which are generic bits of code that are monomorphized over every combination of template parameters that they're used with.
- Virtualization (runtime), classic OO style polymorphism with a VTable (I think this similar to Rust's dyn trait?).
Most exchanges will provide arbitrary depth in way of a pure order feed. From that you can construct your price book, which gives you the levels of depth.
Some just expose the price book, some price book and order feed, some just order feed.
Some anonymize the order feed, some don't (I think typically in equities it's anonymized although not my asset class, but other markets e.g. power need to be de-anonymized and you can see the other buyers and sellers submissions).
However, that data doesn't represent the market price - I would look to constructing the fair price using the trades made on the exchange rather than the outstanding orders. From that you can create different lenses to view the fair price - e.g. volume weighted, time weighted, other types of averages.
> who uses programming as a problem solving tool
Programming can be boring, software development is much more exciting!
All software is solving a problem, same as the software that’s the output of the programming you do.
Those problems might be purely technical (like virtualising different CPU architectures) or focussed on improving the efficiency that others can solve problems (like the software running Stack Exchange) or something completely different.
Solving all of those problems requires more than programming, same as the problems you’re trying to solve need more than programming to fully resolve. Building for maintainability, reliability etc. requires much more than mindlessly programming the software.
But even the programming itself is interesting, especially when solving difficult problems.
> WFH is a rare, once a month kind of thing
years ago for tech in London, way before 2020 for most businesses.
You'd be better off looking outside of London in the commuter belt or further out if you actually wanted to be in an office everyday, where everyone else is also in an office everyday. Or work in a highly regulated industry - defense, something governmental. Or something that requires you to develop against physical hardware.
Codebases that use shared_ptr for everything because they think it ‘solves’ having to manage memory are smelly, it shows that no one has thought about the ownership model. Not to say there isn’t a solid use case for it, but some developers use it by default because they don’t fully understand the trade offs/why it’s important to think about ownership of managed memory.
As an example in our medium to large code base we have two places we use a shared_ptr, completely off the hot path and for long lived objects that aren’t copied more than a few times, but need to share ownership of the pointee - in a similar use case of your use of Arc.
Almost all of the time unique_ptr, with a single owner, is exactly what you want to replace the memory management side of a dynamically allocated object. From there you will pass around raw pointers or references to the object managed by the unique_ptr (for functions that need to ‘borrow’ the object).
Something interesting about shared_ptr is how it relates to weak_ptr, although I’ve never seen weak_ptr in anything out side documentation!
- A bitset index
- Bitflags (e.g. to represent logging levels)
- Feature flags
`bit_count` (aka `popcount`) gives you an efficient way to figure out how many bits within those structures are set and not set.
- Type of generation, looking at the NY data they have a huge hydro generation, obviously the UK grid has much less of that
- Time of day, the Elexon graph is showing you the system price per 30 minute settlement period. You can see that some of those periods weren't anywhere near what you're quoting (there was literally a period where the system price was 0£ in the last 24h). https://www.bmreports.com/bmrs/?q=balancing/systemsellbuypri... makes the volatility much more obvious.
- The weather, it plays a big part in the system price due to price disparity of renewables (commonly wind in the UK grid) compared to oil based and gas generation.
- The interconnect, the UK grid has interconnectors to the EU grid so there's some price impact from that
Typically immutability in C++ is at the interface scope rather than structural, e.g. prefer this
struct MyType {
int x_;
MyType( int x ) : x_( x ) {}
};
void do_the_thing( const MyType& mt ); // mt.x_ = 1; doesn't compile, mt is const
over this struct MyType {
const int x_;
MyType( int x ) : x_( x ) {}
};
void do_the_thing( MyType mt ); // mt.x_ = 1; doesn't compile, x_ is constIf you’re logging so much that you’re saturating a dedicated thread that just reads from a queue in shared memory, does log formatting and writes to a file then you have bigger problems than figuring out how to log because there isn’t going to be a solution that lets you log with a deterministic latency without skipping events or wrapping around your queue.
Of course, the approach scales fine if you can have multiple log files per process - e.g. a normal log file and a message log or transaction log, because you can give each file its own queue and thread.
There's your performance issue, logging doesn't need to be persisted or even processed by your main program thread or any thread/process actually doing the work. Who's logging straight to stdout or directly ::write'ing a file in a production application?
https://gcc.gnu.org/onlinedocs/gcc/Integer-Overflow-Builtins...
All your type traits/SFAINE/concepts/static_asserts happen at compile time, you can’t use them to do any checks on the value of a type, unless it’s a value available at compile time.