This will only be effective if the data you’re working on is smaller than 64 bits. If you’re working with bytes, for example, you might get 8x parallelism.
3,607 karma · joined July 21, 2009
This will only be effective if the data you’re working on is smaller than 64 bits. If you’re working with bytes, for example, you might get 8x parallelism.
I implemented the sentence boundaries, but also thought that the notion of a “phrase” might be useful for such applications: https://github.com/clipperhouse/uax29/tree/master/phrases
It’s optimized for fast reads in exchange for expensive creation.
One can use unsafe for a zero-copy conversion, but now you are breaking the semantics: a string becomes mutable, because its underlying bytes are mutable.
Or! One can often handle strings and bytes interchangeably with generics: https://github.com/clipperhouse/stringish
A grapheme can be multiple codepoints, with modifiers, joiners, etc.
This is true in all languages, it’s a Unicode thing, not a Go thing. Shameless plug, here is a grapheme tokenizer for Go: https://github.com/clipperhouse/uax29/tree/master/graphemes
ReadOnlySpan<T> in C# is great! In my opinion, Go essentially designed in “span” from the start.
I had forgotten, or perhaps never realized, that substrings in C# allocate. The solution was Spans.
Notably, it caused me to realize that Go had “spans” designed in from the start.
The solution is a test that fails when Chrome and Safari have substantially different render times.
(Whether I recommend it, not sure! I did it and then undid it, with suspicion that tests were taking longer due to, perhaps, worse caching of build artifacts.)
I’ve made the mistake of simply splitting on spaces or punctuation, and eventually finding that will do the wrong thing.
Eventually I learned about Unicode text segmentation[0], which solves this pretty well for many languages. I believe the default Lucene tokenizer uses it. I implemented it for Go[1].
[0] https://unicode.org/reports/tr29/ [1] https://github.com/clipperhouse/uax29
I keep coming back to theory-of-the-firm: https://en.wikipedia.org/wiki/Theory_of_the_firm
An industry that accommodates a lot of bullshit work is an industry with high transaction costs. The author hints at that at the end of the article, but leaves it an open question.
Most of the bumbling fails. Some of the bumbling works. And so the most “effective” politicians (and organizations) will survive, even if their success has little to do with deliberate scheming.
I.e., more dollars chasing the same amount of education. A bit more explanation here: https://fee.org/articles/student-loan-subsidies-cause-almost...
NYC (where I live) is wildly inefficient, except at delivering benefits that millions of people consider worthwhile.
I am unclear if this article is quoting the "net effective" rate, where the free month is amortized over a year. Mine is around $2750 effective, while monthly is around $2950.
The problem with mistaking this fear for a fact is that it often leads to an incorrect intervention. (I call this a WMD argument.)
We’d be much better served with much greater caution about what is actually, observably, measurably true. In this case, we’d have to discover the yet-unfound correlation between technical advance and employment rate.