HNHacker News
TopNewBestAskShowJobs

yaantc

1,576 karma · joined August 13, 2012

submissionscomments
yaantc··on Bad NEWS, Emacs
The article is just his author view. Please read the emacs mailing list threads to get the full picture.

GP is correct, and this is quite normal: some people using Emacs master will not follow all the mailing list discussions and commits. So they will notice a change only after it is merged. Nothing wrong with this, and nothing wrong with being unhappy about such a change. What's wrong in my book is the nature of the reaction show in this article (see my comment in reply to @tarsius).

yaantc··on Bad NEWS, Emacs
Thank you @tarsius! Couldn't have said it better.

To add, I have issues with the attitude shown by the blog author.

If you use the development branch, you can't raise hell when there's a breaking change: it's to be expected!

Then it's fine to disagree on some change and discuss this. I read the email thread, and I do not see "arrogance". Just strong disagreement. So yes, converging will take a bit of time... Calling publicly someone "arrogant" for not folding back to your view, and trying to raise the crowd (a good part of whom won't read the thread to make their own opinion) looks like bullying to me.

Saying that his patch to make the change optional has been disregarded, when it was rejected because it not only made the change optional (that would have been OK, and a patch for this asked for) but removed other changes is not honest.

Lastly, pointing out one person to blame when the whole discussion is done with the Emacs maintainers in the loop is also a no-go in my book.

As a close to 30 years Emacs user, thank you to all its contributors! (and to Thierry, as long time Helm user) May their skin by thick, it's unfortunately sometimes needed :-P

yaantc··on Bad NEWS, Emacs
Read the mailing list thread: the patch did much more than making the change optional, it did revert other related changes. That's why it was rejected. Other discussed changes were taken in, and it's not settled yet it seems: the discussion is on-going.

I find the reporting here very one sided and uselessly dramatic. I read the thread and don't see arrogance, just (sometimes strong) differences of opinion. Calling "arrogant" anyone who don't agree and fold to your view, and create drama and draw the crowd against one specific person (the initial change author, OK, but it was done with maintainers in the loop) where the crowd won't check the details is not OK in my book.

yaantc··on Performance: Adventures in Thread-per-Core Async with Redpanda and Seastar
Oh I'm well aware the author didn't invent any of this: I'm currently working on a "VPP lite" [1] implementation for an embedded packet processing application and as part of my "SoTA" review learned about TPC history (in networking at least).

I was just trying to answer your question on why agents could be relevant here, nothing more.

[1] https://fd.io/docs/vpp/master

yaantc··on Performance: Adventures in Thread-per-Core Async with Redpanda and Seastar
Because the article is on developing a "shared nothing" system where sharing is replaced by communications. Actor systems are also about shared nothing and using messages based communications, and they came first historically.

Now to be fair, a lot of actor systems were about correctness and not squeezing the most performance of a platform. Not sharing data means no explicit locking with mutexes (though one can still have deadlocks, as in several actors stuck waiting for each others) and something simpler to analyze: the system is made of communicating state machines. Actors were also used in high performance system like telecom switches (high performance, for their time), but here too correctness was probably the main concern. At the time actors were first used accessing even far memory was not as costly as today, so the cost of synchronization wasn't as bad. But it was already tricky to get right.

Still, actors are perfectly on topic when discussing a shared nothing, message based architecture. They're worth mentioning IMHO, to give some historical perspective. And then one can combine both approaches: a pinned thread per core can dispatch its messages to a set of cooperating actors for example.

yaantc··on RISC-V Board of Directors decision on the Compressed extension
Not really, GP was talking specifically (and exclusively) about the ease of decoding the native ISA.

Yes, modern x86 are wide from the uop cache, but this cache is made of simple, easy to process uop and not native instructions. The native part is still "only" 4 wide. The whole existence of the uop cache is to avoid repetitive native instructions decoding and replace it by dealing with simple uop. That's because this native x86 decoding is expensive and hard to widen, which is exactly GP's point.

You're saying the uop cache makes this a non issue in practice: maybe (cost/efficiency of this cache vs a simpler design? I don't know ;), but that was not the point GP was discussing. The point was only that the C extension is easy to decode, with as practical proof the width of existing implementations supporting it.

yaantc··on Cortex X2: ARM aims high
Maybe it's more than you want, but "Computer Architecture - A Quantitative Approach" from Hennessy and Patterson is a very good reference.

If your main interest is CPU cores even a relatively old version will be enough. From memory recent versions (5 and 6, 6 being the latest) have mostly added content on GPU and data center computing, but the sections on CPU architecture are rather stable.

yaantc··on Static Analysis Tools for C
Clang static analyzer and tidy through CodeChecker [1]. Plus the compiler warnings of course.

This to complement unit and integration tests in simulation, with all sanitizers (UBSAN, ASAN, MSAN, TSAN).

Both are useful. In practice, the runtime tests catch most bugs. But now and then some bug sleeps through and is caught by the Clang static analyzer. It's always impressive to see a display of 30+ steps leading from an initial condition to a bug.

Finding bugs is the the static analyzer really. Clang tidy (a linter) is for code "cleanliness" and avoiding dangerous constructs. But I don't remember tidy finding a real issue, contrary to the static analyzer.

[1] https://codechecker.readthedocs.io/en/latest/

yaantc··on Static Analysis Tools for C
> It's not that Clang is bad, but it has inherent limitations because it only analyzes a single source file and stops analysis when you call a function from another module.

Nowadays that's only the default. But you can enable "cross translation units" [1] support to perform analysis across all the files of an application. It's easier to deploy CTU by using CodeChecker [2].

Also for the Clang static analyzer: make sure the build does use Z3. It should be the case now in most distro (it's the case in Debian stable ;). It will improve the results.

With both CTU and Z3 I'm very happy with the results. Klocwork mostly only reported false alarms after a clean CodeChecker pass.

     [1] https://clang.llvm.org/docs/analyzer/user-docs/CrossTranslationUnit.html
     [2] https://codechecker.readthedocs.io/en/latest/
yaantc··on How LSP could have been better
> clangd is better but can still get things wrong if your build process is complicated your compile_commands.json doesn't perfectly reflect it.

Yes, clangd is accurate but when you have different build variants only one is reflected in the current compilation database. clangd will be perfectly accurate for this current variant, but blind to others.

Would clangd support a merge of several compilation DB into one? If a file appears several times with different options (typically include paths and defines), would clangd handle all variants in parallel or just pick one (first or last)? I haven't tried (yet ;).

It's manageable, and sometimes I'm reverting to pure "text level" changes or searches to work around this.

yaantc··on VSCodium – Open-source binaries of VSCode
Same here ;) But now there's "8 GB and constantly swapping": Eclipse! I logged once on a dev server to test something, it was slow and I checked what was going on. On this 64 GB machine, 8 Eclipse users, each Eclipse using a tad over 8 GB, led to significant swapping. What a world!
yaantc··on Mastering Emacs
Ok, I should have said relatively, or even better "subjectively", small C source base.

The C base supports many platforms and variants, so only a part of it is used in a running Emacs instance. And to the 1+M lines of elisp one typically had many more extension packages. That's the the "relative".

In practice, when I want to change things I don't hit the C layer. Anytime I had to introspect Emacs to change something I was firmly in the lisp world. I don't remember having been limited by something being at the C level, which are things that cannot be changed dynamically. This is really what I meant: in practice to change Emacs the C level (although big in the absolute, and definitely complex) doesn't show up much and hasn't been a limitation at least to me.

yaantc··on Mastering Emacs
This thread is on the "Mastering Emacs" book and from its author, on the topic of mastering Emacs window managements you can read this:

https://www.masteringemacs.org/article/demystifying-emacs-wi...

To be honest I've never played with this part of Emacs. The way I use it I'm fine with the default. But if not it's possible to take control of Emacs internal windows management.

yaantc··on Mastering Emacs
Literally changing the editor code, to do whatever you want. Emacs is a lisp (elisp) environment mostly with a small C based core for performance sensitive primitives. But the user code is at the same "level" as the Emacs core code.

This is very different from most extensible editors, where extensions are sandboxed and have a limited, controlled API to the core editor, which is kept separate.

This makes Emacs very "malleable", and it's likely easier than with most other tool to change it in custom, specific ways. It may seem dangerous, but it comes from a culture where this extensibility is very front and center, with documented design pattern to support it (hooks, customization settings...).

yaantc··on C and C++ prioritize performance over correctness
For static analysis I use CodeChecker, it's a wrapper on top of the Clang static analyzer and Clang tidy (linter). Now also supports cppcheck, but I disabled it (too many false alarms). It's free and open source, and I find it useful. Make sure you use it with a version of LLVM/Clang with support of Microsoft z3 enabled (it's the case in Debian stable, so should be OK in most distros).

For the flags I would start with "-Wall -Werror", then maybe disable some warnings based on the code base / use.

All this assuming a GCC/Clang friendly code base.

yaantc··on Debian celebrates 30 years
If you're comfortable with Linux definitely, I'm using Debian stable both at work (Latitude) and home (Thinkpad) for several years now.

One issue with Debian was the lack of support for proprietary firmware in the stock installer from the Debian web site. A lot of people trying Debian and not aware of this ended up stuck with no WiFi support during installation for example. But this has been fixed, the Debian Bookworm stock disk image includes firmware now.

yaantc··on Cellphone Radiation Is Harmful, but Few Want to Believe It
Often much less in practice: 2G transmission was up to 2W, but 2G has been phased out in many countries. 3G I'm not very knowledgeable but IIRC was up to 800 mW (common case). LTE/4G is 200 mW maximum for sure. 5G is the same except for mmWaves, but that I don't know very well and it's not used much (not at all in Europe, where I live).

There are more powerful band classes, but for fixed wireless access. So for phone use on a modern network, you should now be most of the time at 200 mW max transmit power (same as WiFi BTW).

yaantc··on Ask HN: For advice: I'm a mathematician looking for a plan B outside of academia
You may consider the electronic design automation (EDA) industry. Basically, the tools used to design the ASIC we love.

The industry anticipates a job crisis, with elder people retiring, a shortage a new entrants (not see as sexy) and still a strong need. There's been initiatives around to bring more new blood in. A math PhD with an interest in optimization looks like a good fit.

There aren't so many employers (Cadence, Synopsys an Siemens/Mentor are the 3 bigs), but the domain is extremely technical with an history of pushing the envelope. SAT solving for example has progressed a lot thanks to EDA, and we not benefit from it in software with its SMT extension.

yaantc··on Nvidia’s CEO Is the Uncle of AMD’s CEO
You can find some guides on the topic, for example: https://www.letudiant.fr/etudes/classes-prepa.html.

Very often, when a kid is good enough its school main teacher will mention the possibility of going to a prépa: it's good for a school reputation when they can send kids to prépas.

For sure it's better when the parents are already in the known, but it wasn't the case for me and it was not a show stopper. There's plenty of public information and the importance of "grandes écoles" in France is not a secret.

yaantc··on It's 2023, so of course I'm learning Common Lisp
Calling again and continuing are not the same thing. Sure, with the above trivial example it is. But if the parent function has non idempotent code before calling the missing function (like doing some global change / side effects), then calling again will give a different result than just continuing from the current state.

So is it possible to define the missing function and continue from the same state in Python? I don't think so, but I'm not a heavy Python user (just for small/medium scripts).

yaantc··on Beyond Markdown (2018)
This is what Org mode does. It's still very tied to Emacs, but there's an effort to standardize the Org format. Hopefully this will help its adoption outside of Emacs, it's a nice markup (and a lot more).
yaantc··on Debian 12 'Bookworm' New Features and Release Date
Info on LTS support (5 years): https://wiki.debian.org/LTS Then there is the extended LTS, ELTS (10 years): https://wiki.debian.org/LTS/Extended

Not all the Debian packages are covered by this security support, but there's a tool to check your installation (see the LTS wiki).

yaantc··on U.S. focuses on invigorating ‘chiplets’ to stay cutting-edge in tech
By using several small chips instead of a big monolithic one, it's possible to reduce costs in 2 ways:

1) the yield is better for a small die. For a given density of defect, a big chip will have a higher probability to have a defect than a small one. Basic example: you use 4 chips instead of one, and one defect that would kill the big chip will only kill one of four of the small chips. It's more subtle than this, there are simulators on the web to see the impact of size on cost for those interested;

2) parts of the chip can use cheaper nodes. For example the I/Os not only can use less advanced and cheaper nodes, but those nodes have often better support for analog IPs.

On the flip side, communications that were internal in the big monolithic die now must cross those small dies boundaries. And communications is expensive: you would certainly not want to handle this through a PCB. Instead, more local short range interconnects are used that are much more power efficient than a PCB interconnect (but not as good as in die). These require sophisticated packaging, which adds to the cost. Still for complex chips the net effect is positive, see what AMD did (with Intel now following).

yaantc··on My favorite C compiler flags during development
As I understand it the Clang static analyzer is not intended to be used directly. It's best to use it through CodeChecker, which provide high level commands and do the low level calls to the Clang SA (and also Clang tidy). In particular if you want whole project analysis (CTU, cross translation units analysis) CC will do the all the needed calls in a single "analyze" command.

[1] https://codechecker.readthedocs.io/en/latest/

yaantc··on The end of nuclear power in Germany
Do you believe the possible threat of nuclear waste to future generations, considering our current waste management technologies, to be higher than the impact due to more CO2 generation by running electrical coal plants instead of nuclear ones? That seems a more relevant question to me.

There's no silver bullet, and pointing problems on one side only may help win arguments (maybe?) but is not so helpful when trying to find the "best" option. Which, in real life, usually means the less bad option. Warts everywhere...

yaantc··on Arm wants to charge dramatically more for chip licenses
I'm not talking about chip makers, but about CPU IP providers. Like ARM.

If you say there are not many publicly available products including publicly accessible RISC V cores from those companies I listed compared to ARM, yes this is true. Because they're (for now) mostly used for embedded application, deeply inside complex SoC. Only used by the chip makers developers.

A relatively small company like Andes has their IP in over 10 billions SoC today. A lot of it for things like touch screen controllers, but also more complex things now. But I'm not sure you can find a mass product with an accessible RISC V Andes core.

Still, public access has already happened with Renesas, and it will likely get more common. People tend to focus on the high end, but the most common in high tech is to attack from below. Take a foothold in the low end because you're cheaper (smaller companies with lower costs can deal with lower prices), build a customer base, then grow from this. Until one day you can attack the high end (we're not there yet IMHO). I have access to these companies roadmap and this is what I see them doing, and it shouldn't be surprising.

But this means that at this stage, the majority of RISC V use is invisible to the public. To me this is perfectly normal. Building an ecosystem and maturing a technology like this takes a long time.

IMHO we're already at a stage where ARM has real competition in the low and mid end, particularly where there's little dependency to the ISA (embedded, or simple uC programmed in C anyway with no 3rd party binaries included, like Renesas). I think ARM is still safe at the high end though, and it will take some time for RISC V to get there. Chip making is hard, particularly at the high end.

yaantc··on Arm wants to charge dramatically more for chip licenses
In the low and mid end there's plenty of competition. Andes, SiFive, Cortus, BA Semi and probably more. Many of them used to offer their own proprietary ISA and have moved to RISC V to leverage the ecosystem (software and tools). Instead of a semi frozen old GCC version their customers can now use a recent compiler directly from upstream for example.

Renesas has started replacing ARM core by RISC V in their new micro-controllers. They still offer their ARM products and will do for a long time, but move new ones to RISC V.

ARM's advantage is at the high end. For a high performance core it's key to work closely with TSMC and Samsung on the most advanced nodes, to tune the design to what leading edge processes can do. This is very labor and capital intensive. It definitely gives ARM an advantage against smaller IP providers at the high end. But at the low and mid end? Nope. It's where ARM is easier to replace.

yaantc··on Show HN: GPT Repo Loader – load entire code repos into GPT prompts
Am I missing something? From what I understood from Wolfram description of GPT and GPT in 60 lines of Python, a GPT model's only memory is the input buffer. So 4k token for GPT3, some more but still limited for GPT4.

To summarize the GPT inference process as I understood it, with GPT3 as example:

1) the input buffer is made of 4k token. There are about 50k token. So the input is a vector of token ids. We can see it as a point in a high dimensional space;

2) The core neural network is a pure function: for such an input point, it will return an output vector as large as there are token. So here, a 50k element vector, where each entry is the probability that the associated token is the next element.

The very important thing here is that the whole neural network is a pure function: same input, same output. With immensely large super fast memory this function could be implemented as a look-up table, from an input point (buffer) to an output probability vector. No memory, no side effect here.

3) The probability vector is fed into a "next token" function. It doesn't just take the highest probability token (boring result), but use a "temperature" to randomize a bit, while using the output probabilities;

4) The next token chosen is inserted into the input buffer, keeping the same total number of token. Go back to (1) until a "stop" token is selected at (3).

So in effect, the whole process is a function from a point to a point. "point" here is the buffer seen as a (high dimensional) vector, so a point in a high dimension space. The generation process is in effect a walk in this "buffer space". Prompting puts the model into some part of the state, with some semantic relation to the prompt semantic content (that's the magic part). Then generation is a walk in this space, with a purely deterministic part (2) and a bit of randomization (3) to make the walk trajectory (and its meaning, which is what we care about) more interesting to us.

So if this is correct, there is no point in injecting a lot of data into a GPT model: the output is defined by the input buffer size. Just input the last 4k token (for GPT3, more for GPT4) and you're done: everything else would have disappeared. So here, just input the last 4k token of a repo and save some money ;)

To avoid this limitation, one would have to summarize the previous input, and make this summary part of the current input buffer. This is what chaining is all about if I understood correctly. But I don't see chaining here.

Sooo... Am I missing something? Or is the author of this script the one missing something? I don't mind it either way, but I'd appreciate some clarification from knowledgeable people ;)

Thanks

yaantc··on Use GNU Emacs
M-x untabify ? ;)
yaantc··on Use GNU Emacs
It's not so simple.

With UNIX, the integration platform is the shell, hooking simple specialized tools through pipes and the shell language.

With Emacs, the integration platform is a Lisp environment, hooking specialized Lisp tools and Unix tools through Emacs buffers (or buffers regions) using the Elisp language.

So it's an extra layer, but if you look at it in this way you can see similarities: an environment that makes it easy to compose elementary functions into an integrated whole.

In both cases, it's most suited for people who are ready and willing to build their own specialized environment on top of a powerful platform. Although many users don't and stick to the basics too.

← PreviousPage 2 of 14Next →