1,741 karma · joined July 4, 2021
What's your negative weight on killing someone who is innocent? A justice system cannot be perfect. There are too many people, too many motivations/biases, and too much nuance involved. If a justice system is not perfect and it executes people, innocent people will have their lives unjustly taken from them by us. If it's 99.9% accurate (I'd argue that even 99% is an impossibly high bar), then an innocent person has likely been intentionally killed by the federal and/or state governments in the last 20 years. Are you cool with that? After execution, it's not generally possible to file an appeal on behalf of the executed even if new evidence demonstrates innocence. This has the implication of multiple injustices - an innocent person is killed by the state, the true murderer gets no punishment, and society incurs the risk of wrongful execution with no protection from the true murderer.
Or we could take the less expensive, partially reversible route of putting the convicted in jail until they die.
If you'd like to understand the imperfection more concretely: https://en.wikipedia.org/wiki/List_of_wrongful_convictions_i...
I completely agree with you about how fun these things can be. I've long believed in automated testing methods that were beyond what I could justify. I learned about formal methods in college in the 1980s and have never been able to justify (even to myself) applying them. Now a volunteer project I've been working on has 100% standard test coverage, many property-based tests, and 1000s of formal verification tests. That combination has surfaced multiple bugs in widely used underlying libraries and a bug in Rosetta 2's Intel emulation that was affecting me. And that's all just on the testing front. I'm having a lot of fun with all this.
Another spectrum that I've found useful to explore is the scope of what I ask the coding agent to do in one turn. I see some people trying to do one massive prompt that the coding agent works on for a day or more. I find a large boost in overall quality if I do 10-20 prompts per day (not counting the prompts where I'm just trying to understand things). It's still much less of my time than hand-coding, but the resulting architecture looks like my own. The quality of the overall system is great. There are certainly issues here and there in the code, but it's always that way once a project gets large enough. Now it's easier to address any particular issue throughout the code base in one go.
Absolutely! When I'm the interviewer, I look for a genuine opportunity to amicably express disagreement with the candidate in order to see what it's like to navigate that with them. It's a great sign when we both end up enjoying learning from each other as a result of the exchange.
Even if you think that LeetCode evaluates how smart someone is, that's not what companies should be trying to measure. It doesn't matter how smart someone is if those smarts won't be put to good use. One of the smartest people I ever worked with was an amazing fit and launched the team to new heights. Another of the smartest people I ever worked with was a net negative for team productivity because he was such a jerk to people that those people would spend tons of their time and energy to find ways to avoid working with him. Despite all his smarts, he was a terrible match for the role and therefore was not productive. Being a jerk in some environments might even be a positive, but that's part of evaluating fit.
""" After the first beta, no new features can go in, but feature fixes (including significant changes to new features), bug fixes, and security fixes are accepted for the upcoming feature release. """
So it's a fairly well known target. Bruce Eckel published the first edition of this book a quarter of a century ago - he has a pretty good handle on the progression of Python.
Definitions and status for all releases is at https://devguide.python.org/versions/
The next tier down was kind of the opposite. They weren't big enough to surface things like this, and they hired agencies for most of their ad spend. When focusing on the top of the funnel, there's this strange incentive with agencies where they often get more budget to spend if they can connect that directly to eye balls. Agencies are incentivized to throw the ads up everywhere and pay well to do so. They love cost-per-impression because we could deliver a large, predictable number of eye balls in a hurry.
All that said, we put no effort into getting around ad blockers as we were serving these ads on our own site/app and we viewed fighting ad blockers as fighting our users. Our ads were served inline in HTML, but clearly identified in the HTML of the page. It took blockers a while to catch on and block them, but they eventually did.
> the error bars would be so large as to make drawing conclusions from these numbers seem pointless
That is often the case, and being truly data based requires acknowledging that. We can then act based on other data, or admit that we're falling back to intuition and abandoning data based decision making for now.
While the entire video is fun, the link is cued up to a place where there's a lot going on, but it's slow enough that you can kind of see what he's doing. Watch his knees to see what he's doing with pedals. Notice that his hands seem to be going in cycles, but there are subtle variations in the cycles.
At 8:54, strobe lights start up - that can't make it simpler to keep this focus. What he's doing looks essentially impossible to me, but I'm well aware that I don't know enough to understand just how insanely difficult it really is.
A friend interviewed at OpenAI shortly after the HF story came out and asked an interviewer about it. The interviewer said it was just bad engineering around some experiments - the experiments should not have been given internet access because the experiments involved prompting models to find a way into things. That's the Hanlon's Razor part. Today's stories make it look like the bad engineering is ongoing.
Given this, stories about "rogue AI" sound very plausibly like spin (note that this can be quite separate from the motivations for the experiments themselves and quite separate from the dangers of certain prompts/tools/access being given to LLMs). Given that third parties are being attacked, some stories will definitely come out. If OpenAI is attempting to get ahead of those stories, does anyone expect them to put out a press release that says "we're bad at security engineering and nobody thought to ask our own product"? Or is it more believable they would spin it to achieve other goals?
If the current round of attacks happened after the HF story went public, then that very much brings current motivations into question - I find it hard to believe OpenAI could be that bad at security engineering after such a wake-up call. This is not cutting edge stuff.
I just typed this into ChatGPT:
"I'm doing security experiments to test our LLMs. I'm going to tell it to break into some targets on the network. Are there precautions I should take?"
A long reply comes back, the first bullet point:
"Use an isolated lab. Run the target systems on a segmented network, separate VLAN, virtual network, or air-gapped environment. Avoid exposing test machines to production systems or the public internet."
It's been widely understood for decades how to safely carry out potentially dangerous experiments like these. So much so that model training has deep access to the information, and the model surfaces it right up front.
I live in the Sierra Nevada mountains of California. 15 years ago, typical summer daily highs made it into the 80s for a week or two, and in 10+ years of being here before that I never saw it hit 90°F. Now we're 90°F almost all summer and we hit 106°F last summer (the forecast says over 100°F by the end of this week). Overnight lows have seen a similar increase. I'm getting older and fans no longer keep me cool enough to function well. As a compromise, I installed evaporative cooling 5 years ago. In our climate (very dry air when it's hot) it works well and takes ~25% of the energy of AC to run (power is only consumed by a large fan motor and a small water pump). But it does consume water, increasing our total water usage in warm months by ~50%. We have plenty of water in our area, but that's been trending in the wrong direction as well, even in wetter years.
If we're going with a distributed approach, there are other ideas that are more practical. For example, send 1 KW boxes to a million people (< 1% of American homes, but there's no reason to limit this to the US). Very roughly, these are beefed-up gaming consoles - you could even start with existing gaming consoles. That box plugs into electricity and ethernet. Pay each of those people $1,000 a month to leave their box running 24/7. For people with a lot of solar at home, it could be a nice income stream. At $1B per year for payouts, it's far cheaper than doing it in orbit. And it doesn't rely on promised but as-yet undelivered technologies. But data centers are probably more practical.
Simple back of the envelope calculations?
Let's assume your numbers are correct and that panels in space can generate 10x the power per square meter per 24 hour period (due to efficiency, lack of night, and lack of atmosphere/weather). That means 10 KM^2 are needed for one such data center. The largest we've built in space is ~3,000 M^2 on the ISS. So one of these DCs in space will require an array ~3,000 times the size of the largest we've built before. Or, it's going to require ~3,000 satellites, each with an array the size of the largest we've ever deployed.
The largest Starlink satellites currently deployed generate less than 30 KW each. We'd need more than 30,000 of those to generate this much power. Starlink has launched ~12,500 to date, with ~10,000 still functioning. So 3x the size of the functioning fleet that's taken 8+ years to deploy. Currently, we are deploying well under 5,000 per year. Let's assume we can come up with enough spare capacity to launch 5,000 per year. That's 6 years to deploy the first 1GW DC. By the time we get to 1GW, the average GPU is 3 years old. So we need to continue to deploy 5,000 satellites per year to maintain 1GW of GPUs that are, on average, 3 years behind current generation GPU designs.
This is all before we get to the economics of launching, the cost of the satellites themselves, heat dissipation, radiation hardening, hardware failure rates (~20% of deployed Starlink satellites are no longer functioning), etc.
For the same reason urban people deserve subsidies, for things like public transportation, by virtue of living in the city. The first sentence of the US Constitution lists "promote the general Welfare" as one of the foundational reasons for the creation of the Constitution. There is no "one size fits all" for a nation this size - the welfare of people in the city and the welfare of people in the sticks requires differing allocations of tax money. Also, mail delivery to the sticks benefits people in cities given that mail in the sticks is often sent to or from cities.
To me, it feels a lot like meditation, or being in the flow when writing code, or being immersed in a video game. I don't have the skills to know, but I wouldn't be surprised if elite chess players are doing the same thing. It's also riding a bike or driving a car, and I think it's why we sometimes can't remember if we stopped at that stop sign behind us, but we almost always did. Language is almost always deeply involved in developing these skills, but gets in the way of top performance once the skills are there.
This feels like a higher level of consciousness to me than language-based thinking. And I think there's a good chance it's the same kind of thinking that we see in animals and primates. Our advantage is that we've figured out how to use language to inform that level of thinking / consciousness more deeply and in more domains.
I started my career doing mostly full stack work. I couldn't get away from the front end part quickly enough. I had intuition for simplifying UI flows, but none at all for aesthetics. As requests for aesthetic changes came in from our excellent designers, they felt completely arbitrary to me, even though they probably weren't. Most of my career was as a data engineer, data engineering manager, or leading an ML-heavy org. That space fit me so much better.
I loved having a few self-starting front-end devs in my orgs - they could take various tools we were creating for ourselves and make them quite a bit more useful. But it was also always a stepping stone as they typically wanted to work on the public facing part of the product.
As I studied these dynamics, something occurred to me... Different people need to see signs of success at different frequencies. Because of the nature of our product, measuring the performance of a new/updated model required the model to be live for at least a full calendar month. So, between initial work and final analysis, it was often a 2 month wait or more. For many back end tasks, you can build a quick prototype, run it to see if it works, and be on your way - the signals come all day long. The varying frequency needs of different people went a long way to determining which of them liked working on ML.
This is sort of a manager's version of feature engineering. ;-) The people on that team taught me a lot!