602 karma · joined December 2, 2024
FWIW I think most people don't belong to either group exclusively. I often care mostly about the outcome (e.g. at work or when i make small tools for myself), but occasionally i care mostly about the process (e.g. when learning a new programming language or solving problems akin to advent of code)
> "System 1" is fast, instinctive and emotional
this implies more than just "fast", which is precisely why i don't like its present usage
I was curious about this so I skimmed the paper [0]:
> SalesRLAgent achieved 96.7% accuracy, outperforming the best commercial alternative by 23.7 percentage points and the best LLM approach by 34.7 percentage points.
For a fuzzy natural language task like this, this magnitude of improvement should already set off alarm bells (Though i admit I'm not even sure what accuracy is even measured here, and the paper doesn't help either). Also, "best LLM" here refers to GPT-4 (at the time of upload, the public already had access to GPT-o3 and). I would have loved to contextualize the performance by looking at model size, but the paper is frustratingly devoid of detail in that regard:
> The core of SalesRLAgent is a reinforcement learning architecture consisting of: • A state encoder network that processes Azure OpenAI embeddings and features • A policy network that estimates conversion probability based on the current state • A value network that estimates the expected cumulative reward • A meta-learning module that assesses prediction confi dence
Also:
> Beyond technical metrics, we evaluated SalesRLAgent in real-world sales environments through A/B testing. [...] After 90 days across 217 representatives and 12,433 con versations, we observed: • 43.2% increase in conversion rate for the test group
This would be a pretty huge result but the fact that this is just shoved into a single paragrpah with no further discussion on methodology, baselines and setup makes me very suspicious.
Also I feel like the obvious way to read the very first sentence is that GLM is a language model
> As we develop GLM, the model sometimes exhibits capabilities that surprise us
> The [k-server] problem’s definition is simple: There are k servers located at points of a metric space. At each time step, a request arrives at a point of the metric space. An online algorithm must serve the request immediately by moving a server to the requested location, without knowledge of future requests. The goal is to minimize the total distance traveled by servers.
> The k-server conjecture states that a deterministic online algorithm can achieve competitive ratio k on every metric space.
I only had to look up what "competitive" means in this context, and wikipedia [0] had this to say about it:
> An algorithm is competitive if its competitive ratio—the ratio between its performance and the offline algorithm's performance—is bounded.
The ratio by which this performance is bounded for a k-competitive algorithm is k (plus some constant) [1]. We can consider the analogy of k support technicians ("servers) located in different (physical) locations ("in metric space"): The conjecture/theorem states that in any metric space (Not necessarily two- or three-dimensional), there exists an online algorithm that results in travelled distances of no more than roughly k times that of the optimal distance if all requests were known in advance.
[0] https://en.wikipedia.org/wiki/Competitive_analysis_(online_a...
[1] https://www14.in.tum.de/personen/albers/papers/brics.pdf Section 1.1
which links to: https://archive.org/details/s9notesqueries03londuoft/page/12...
which is in reference to the original proquiritations here: https://archive.org/details/worksofsirthomas00mait/page/416/...
i had also never heard of this before today and wonder if people had even seriously tried to decipher this at all?
For false negatives, that depends on the type of article at hand. If an error leads to me spending 2 minutes reading an article that ended up being not worth reading I'd be fine with an order of magnitude more than if false positives happen on articles that take 30 minutes to go through. Maybe on the order of 1-3% overall?
Not quite but close enough :D https://www.youtube.com/watch?v=DWzWwx8T3nE
Are there any regular typing keyboards with velocity sensitive keys? I genuinely don't know, it's just the first thing I'd compare this to