5,507 karma · joined October 29, 2009
Enthusiastic geek https://idea.popcount.org
https://github.com/majekIn more sophisticated configs adding / removing IP's or TLS certs requires restarting server, configuring applications. This gets out of hand quickly. Like what if your server has primary IP removed, because the IP space is recycled.
At CF all these things were just a row in database, and systems were created to project it down to http server config, network card setting, BGP configurations, etc. All this is fully automated.
So an action like "adding an IP block" is super simple. This is unique. AFAIK everyone else in the industry, back in 2012, was treating IP's and TLS more like hardware. Like a disk. You install it once and it stays there for the lifetime of the server.
I liked the article. It's close to my adverntures with containers. I think the invention of docker is indeed mostly packaging: Dockerfile (functional) is pretty neat, docker hub (addressing a container) is awesome, and the ENTRYPOINT in Dockerfile is great, it distinguishes Docker from .deb.
But indeed, beyond Dockerfile things are bleak. Docker compose never rose to my expectations. To get serious things you need load blancer, storage, addressing, and these are beyond traditional containers scope.
Back in Jan I wrote a piece about how to actually use MPTCP
https://blog.cloudflare.com/multi-path-tcp-revolutionizing-c...
But plenty has changed since then. It seems all my complains about the API are now addressed. Maybe it's a good time to actually run with MPTCP again :)
In my private affairs, I realised I need MPTCP less, since I started using tailscale. My SSH sessions tend to last longer when going over it.
I'm not following. I thought the whole issue is that users _do not_ want to use microsoft account locally and that microsoft fights that.
We are a young startup building modern software for European defence. Looking for generalists, especially with experience in: Web backend (FastAPI, Python), Audio processing (including Voice To Text models), Frontend (modern JS frameworks)
https://notion.lysk.ai/engineer
Email me: marek [@] lysk.ai
However they rely on Trampolines: https://gcc.gnu.org/onlinedocs/gccint/Trampolines.html
And trampolines need executable stack:
> The use of trampolines requires an executable stack, which is a security risk. To avoid this problem, GCC also supports another strategy: using descriptors for nested functions. Under this model, taking the address of a nested function results in a pointer to a non-executable function descriptor object. Initializing the static chain from the descriptor is handled at indirect call sites.
So, if I understand it right, instead trampoline on executable stack, the pointer to function and data is pushed into the "descriptor", and then there is an indirect call to this. I guess better than exec stack, but still...
I guess this post assumes the need to use all the gpu resources from within a single block.
On one hand it's trivial - counts (ticks) per second. However, this (of course!) can be very spiky. I ended up using pretty simple EWMA to smooth the results for user interaction. Anything really works, short decay is fine.
Then the really fun bit, was trying it with more serious radiation source, and guess what.... interrupt per tick, is... really bad! I was easily able to overwhelm the arduino, too many interrupts. Fun project to understand interrupt masking.
These were usually quite short notes. Like I did this and that. Or here are instructions I've run. Pretty much a "lab notebook".
Again, this was often useful to people. Folks new what I was working on at that time. Over time more people adapted this style of public note taking (within company), but I have the impression I was more consistent than others.
I did it for many reasons. I genuinely forget things. It was always nice to see comments (like James complaining that I shouldn't do `dpkg -i x.deb`, and instead finally install the package repo in apt-sources!). So again - searchability (a form of documentation), reach (getting feedback), and work log (for planning).
The format was also nice - short, without specific audience in mind. Zero drama, zero effort. Plentiful and low quality. Because it was "blog" and not "pages", this meant it was obvious when the note was written and nobody expected old blogs to be up to date. The lower the friction in note taking - the better.
Finally, I was working with an SRE team, which was tightly knit and communicated very often over daily checkins (which I wasn't invited to) and other informal channels. And I worked from home a bit. This meant I had to find an asynchronous way to communicate with the team. My personal blogs worked nicely. I highly recommend this style to anyone working remotely.
Although I must admit the personal "lab notebook" wiki thing is not scaleable. It's impossible to "follow" more than a handful of people.
I kept statistics. It's fairly obvious when I was most productive. Here it is, my internal blogs on the wiki per quarter:
2013-Q4 | 0 |
2014-Q1 | 0 |
2014-Q2 | 2 | ##
2014-Q3 | 3 | ###
2014-Q4 | 2 | ##
2015-Q1 | 14 | ##############
2015-Q2 | 15 | ###############
2015-Q3 | 57 | ###################################################
2015-Q4 | 60 | ######################################################
2016-Q1 | 70 | ######################################################################
2016-Q2 | 71 | #######################################################################
2016-Q3 | 17 | #################
2016-Q4 | 23 | #######################
2017-Q1 | 13 | #############
2017-Q2 | 22 | ######################
2017-Q3 | 16 | ################
2017-Q4 | 25 | #########################
2018-Q1 | 12 | ############
2018-Q2 | 9 | #########
2018-Q3 | 23 | #######################
2018-Q4 | 14 | ##############
2019-Q1 | 14 | ##############
2019-Q2 | 20 | ####################
2019-Q3 | 18 | ##################
2019-Q4 | 36 | ######################################
2020-Q1 | 19 | ###################
2020-Q2 | 13 | #############
2020-Q3 | 4 | ####
2020-Q4 | 3 | ###
2021-Q1 | 4 | ####
2021-Q2 | 8 | ########
2021-Q3 | 3 | ###
2021-Q4 | 21 | #####################
2022-Q1 | 4 | ####
2022-Q2 | 12 | ############
2022-Q3 | 4 | ####
2022-Q4 | 28 | ############################
2023-Q1 | 14 | ##############
2023-Q2 | 16 | ################
2023-Q3 | 10 | ##########
2023-Q4 | 6 | ######
2024-Q1 | 12 | ############
2024-Q2 | 22 | ######################
2024-Q3 | 0 |
2024-Q4 | 2 | ##
2025-Q1 | 20 | ####################
2025-Q2 | 3 | ###
Total 784 over 11 years.In my case - indeed the name is a historical baggage, I'm not arguing for or against it.
Indeed we had regularly situations that we had to pull in experts from other rooms, to discuss specific topics (like TCP), so we should have forwarded the conversation at the start.
But I don't think this should be categorical. There is value in non-experts responding faster (the channel had good reach) by your non-expert colleagues than waiting longer for the experts on the other continent to wake up.
Maybe there should be an option to... move conversation threads across channels?
I think there is place for both - unstructured conversations, and structured ones. What I don't like about managerial approach, is that many managers want to shape, constrain, control communication. This is not how I work. I value personal connections, I value personal expertise and curiosity. I dislike non-human touch.
"You should ask in the channel XYZ" is a dry and discouraging answer.
"Hey, Mat worked on it a while ago, let's summon him here, but he's in east coast so he's not at work yet, give him 2h" is a way better one.
I know that concentrating knowledge / ownership at a person is not always good, but perhaps a better way to manage this is to... hire someone else who is competent or make other people more vocal.
And yes, I don't like managers trying to shape communication patterns.
I quickly realised that these conversations had value outside the two of us - pretty much everyone else onboarded had similar questions. Some subjects were about pure onboarding friction, some were about workflows most folks didn't know existed, some were about theoretical concepts.
So I moved the questions to a public (within company) channel, and called it "Marek's Bitching" - because this is what it was. Pretty much me complaining and moaning and asking annoying questions. I invited more London folks (Zygis), and before I knew half of the company joined it.
It had tremendous value. It captured all the things that didn't have real place in the other places in the company, from technical novelties, through discussions that were escaping structure - we suspected intel firmware bugs, but that was outside of any specific team at the time.
Then the channel was renamed to something more palatable - "Marek's technical corner" and it had a clear place in the technical company culture for more than a decade.
So yes, it's important to have a place to ramble, and it's important to have "your own channel" where folks have less friction and stigma to ask stupid questions and complain. Personal channels might be overkill, but a per-team or per-location "rambling/bitching" channel is a good idea.
- they indeed seem to have trained on movies/subtitles
- you absolutely positively must use Voice Activity Detection (VAD) in front of whisper
https://docs.rs/subsecond/0.7.0-alpha.1/subsecond/index.html...
"In practice, frameworks that implement subsecond patching properly will throw out the old state and thus you should never witness a segfault due to misalignment or size changes. Frameworks are encouraged to aggressively dispose of old state that might cause size and alignment changes."
I don't think "throwing out the old state" is a sensible recommendation. "re-instancing" is called "update/upgrade/downgede of internal state" in OTP:https://www.erlang.org/docs/24/man/gen_server#Module:code_ch...
Perhaps I'm missing something, maybe subsecond is a good tool for toy apps or for a developer workflow. But for anything serious, I'd think that managing layout of structs is a primary concern.
I found it easy to start. Then there was a pretty nice learning curve to get to warps, SM's and basic concepts. Then I was able to dig deeper into the integer opcodes, which was super cool. I was able to optimize the compute part pretty well, without much roadblocks.
However, getting memory loads perfect and then getting closer to hw (warp groups, divergence, the L2 cache split thing, scheduling), was pretty hard.
I'd say CUDA is pretty nice/fun to start with, and it's possible to get quite far for a novice programmer. However getting deeper and achieving real advantage over CPU is hard.
Additionally there is a problem with Nvidia segmenting the market - some opcodes are present in _old_ gpu's (CUDA arch is _not_ forwards compatible). Some opcodes are reserved to "AI" chips (like H100). So, to get code that is fast on both H100 and RTX5090 is super hard. Add to that a fact that each card has different SM count and memory capacity and bandwidth... and you end up with an impossible compatibility matrix.
TLDR: Beginnings are nice and fun. You can get quite far on the optimizing compute part. But getting compatibility for differnt chips and memory access is hard. When you start, chose specific problem, specific chip, specific instruction set.
$ strace -e trace=clone -e fault=clone:error=EAGAIN
random link: https://medium.com/@manav503/using-strace-to-perform-fault-i...It turns out - there is now some reasonable tooling to understand verifier!
For stack problems, clang accepts `-s` which prints stack requirement per function, like so:
** stack usage by function **
ebpf/ebpf_aes128.c:180 AES_ECB_encrypt 32 static
ebpf/ebpf_sha256.c:34 sha256_calc_chunk 64 static
ebpf/ebpf_sha256.c:123 sha256_hmac 40 static
ebpf/ebpf_quic.c:86 compute_hp_mask 24 static
ebpf/ebpf_quic.c:242 decrypt_quic 16 static
ebpf/ebpf_quic.c:193 _do_decrypt_quic_loop 16 static
And for instruction count, I was able to feed the logs from verbose verifier (during loading) into code-coverage tooling, and count stuff up. The reuseorg prog takes 100k verifier instruction count/paths: ** verifier instruction count **
udpgrm_reuseport_prog processed 103486 insns
udpgrm_setsockopt processed 9260 insns
udpgrm_getsockopt processed 4215 insns
udpgrm_bpf_bind6 processed 75 insns
While 100k is lower than 1m instr count limit, it's still a lot. And reordering some loop or introducing some "if" often makes that count baloon. Remember that verifier instructions is not real instructions during run. Rather it's a pessimistic interpretation of how many instr max under pessimistic conditions could possibly be run. I don't think having an actual run of that max is even practically possible. Think about it as upper bound of static analysis.Anyway - with stack and instr statistics it's way easier to make sense of verifier problems.