415 karma · joined June 20, 2024
That actually made me LOL
And when this results in actuators executing some bad actions they scream in horror "AI went rogue! It escaped the containment!!! It's going to kill us all!!!"
Go fix your software before you let it do stuff online or IRL. It's not "Terminator", it's just bad QC.
Go to the most popular LLM (ie Claude). Give it the job description. Give it your resume. Ask it for percentage match, it will give you a number. Then ask for advice on improving your resume, you can either ask to rework the resume by itself (if you're in a hurry) or you can work with the thing to modify the sections/entries yourself, one by one, if you have time. With each modification you will get better and better percentage match. Ideally you need to get to high 80% or maybe even 90%.
The idea is that automated ranking tools used by recruiters/HR are ultimately use the very same LLM, so you are improving your ranking assigned to your resume by those ranking systems.
I have to add that now this advice is mostly useless as everybody's using the same technique so you won't be standing up against others but merely be on the same level.
Add to it the fact that many job descriptions are inaccurate and sometimes downright misleading, and hiring managers use their own criteria, and you will understand why connections/referrals is probably the only way to get hired now.
How computer program arrives at the result is utterly irrelevant, through explicitly written instructions or through running inference on pre-trained neural network. What matters is that it does not have agency. Its creators and operators do. So whatever the software does they are responsible for, both good (summarizing my emails for the week) and bad (gaining unauthorized access and destroying data). There's no need for new anything, it's all covered in existing legal frameworks (including presence or absence or intent).
> The bit where, even without anyone bringing up "p(doom)", a system which has the means to hack arbitrary other machines, and which appears to be > motivated to do so by accidental mis-phrasing of prompts, can obviously cause damages exceeding the USA's GDP, let alone whatever public liability > insurance the company happens to have.
Yes, absolutely, which means that building actuators that convert the output from probability-based, black box, non-deterministic systems, that are known to produce unexpected output, into actions in the real world is absolutely horrendous idea. Bizarre even.
What's your point?
What exactly requires "new regulatory framework" here? You running your software resulted in illegal actions, you are to be held liable within existing laws and regulations.
Make that make sense?
It was really, really, really sloppy approach. And all could have been corrected just by executing go-around literally at any point in time up to 10 sec before he did. Both pilots still would have their careers, and those 5 people would still be alive.
Still, the PM notified PF that they were going too fast many times during the approach. All warnings were ignored.
The pilot flying appears to have been dead set on landing on the first try no matter the cost and the first officer did not intervene as he should have.
They got incredibly lucky and both survived, however likely will never be behind the controls ever again.
- Pilot Flying (PF) was the Captain, 55 yo (presumably) more experienced pilot, who only got certified to fly that type (B767) in May
- Pilot Monitoring (PM) was a 38 yo first officer, also not very experienced on B767
- PF made a number of mistakes, going too high, too fast (partially because he was descending too fast), not catching the glidescope on time
- PM was concerned that the approach/descent was abnormal but kept his mouth shut as the more experienced captain was at the controls (Edit: apparently he did speak up but PF did not acknowledge)
- PM lowered the flaps much later than required (PF orders flaps lowered and PM executes, but flaps can only be lowered when aircraft goes below certain speed) and the flaps were lowered all the way as the plane crossed the runway threshold, which likely caused plane to "float" and overshoot even further
- Plane touched down too late, and left wheel even later, which likely precluded some braking mechanisms from activating (or the PF forgot to since was busy trying to put the thing down)
- PM finally said "too fast, go around" but PF ignored him (or likely did not hear)
- In these situations PM has the right (and obligation) to take over the controls and start go around by themselves, but he didn't, presumably because the Cap was driving
- PF finally panicked and started the go around procedure, bit realized it was too late almost as the front wheel exited the runway, idled the engines and reapplied the brakes.
So, yeah...
I found it lacking details. All these things do is split the input into tokens, encode them, feed to the input, run the inference, collect the output, decode, convert to tokens, assemble the final text, and output (leaving aside multi-modal capabilities for now). What exactly can be exploited here? Aside from usual things like buffer overruns etc.
It does mention that vLLM at some point ran eval() on final output (?) which was confusing, why would it do that? It's an inference engine, not an agent.
So interesting topic, but lacks details.
Edit: many (all?) of them have http endpoints, so that obviously can be exploited, but I don't think it would qualify as exploiting inference, it's just hacking the http service.
Thank you in advance!
>Similarly, software engineering is about applying the foundations that mathematics, electrical engineering or computer science brought in a similar scientific way.
Exactly my point, that's why most of IT professionals should be called "software engineers" and not "computer scientists", just like electrical engineering is not "electrical science" or "physics" even though they it's obviously applied physics.
A graduate with physics degree will not work as electrical engineer, what's with all the "computer science" degrees required to get into software engineering?
What we call things is important because that's how we convey meaning/definition of things. Because when we say "computer scientists" people will start comparing them to other scientists which is completely inaccurate.
>Engineers don't traditionally even "build" anything and there are both upper-case and lower-case engineers in title and legal designation.
Not sure what the difference is between Engineer and engineer (again, definitions are important), but I would be ok with equating programmers with other trades...people, like electricians, car mechanics, carpenters, etc. And then higher-end IT professionals (infra/data/application architects) would be akin to engineers, designing, but not building (for the most part).
But neither category would still be scientists.
"Computer scientists" do exist, they create theoretical underpinnings of the hardware and software that millions of engineers then use to create things, but majority of people working in IT are engineers.
Now, don't get me started on "data scientists"...
Still, this would have been hilarious if it weren't so sad...
I haven't seen anyone make an argument they are as good as SotA (OpenAI, Anthropic). It's just they are approaching state where they are "as good" for some _limited_ set of use cases. Which will allow us to resolve 2 primary issues with these SotA models: privacy and vendor lock-in. Plus, they're very useful for education purposes, you get to explore what things looks like under the hood, play with various models, tools, maybe put something simple together yourself.
You get Macbook - great. You got gaming rig with a decent GPU - great (set it up as a dedicated server that you connect to through simple REST).
What exactly is wrong with any of that?
Given that US legal system is precedent-base that... changes things.
"Open weights model" means the developer made the model available for everyone for free. You can download it from huggingface.co for example and do whatever you want with it.
Why "open weights" and not "open source"? Because the "source code" for LLM would include things like training data, training methodologies and tools, so that you can do the training and produce the model (files) yourself. That would be like compiling from source code. Which is not done with these models, it's company's know-how, they only share the end result.
It's more analogous to "freeware" which is what we traditionally call freely distributed binary executable files. But people started calling them "open weights" instead and the term stuck.
Developing these things is NOT free, there's a lot of labor, hardware, compute/memory/storage/network that goes into that. Who's paying for all this? Chinese govt? Developers themselves? What's the revenue model here?
I absolutely LOVE ability to either run them locally or access inference providers on the cheap, but having a hard time understanding the financial side of this.
There's many providers that run open weights models and give you access. Many decent open weights models cannot be run on consumer-grade hardware (DeepSeek, GLM, many others).
Should be fun.
Edit: clarification
They are not "spying", they are _legally_ using various data _legally_ collected on their current and prospective customers to set the pricing. Do you ever read T&C of any online services you're using? So how is this illegal?
>Can you imagine the mayhem if companies just straight up knew salary information for all of their customers?
I filled in FAFSA and CSS Profile for multiple colleges my child applied for last year. So yes, that's exactly how some industries work. Sucks for me, the customer. Not illegal.
Obviously as a hiring manager you're looking for a hard working individual with a number of successfully completed projects and glowing referrals from multiple places of employment, but you're also looking for a person with expertise in particular technologies/industries/whatever other areas of expertise. To a large extent requirements for each role are unique, however many do have some overlap.
So being rejected from one position might simply mean there's a misalignment between what the company is looking for and what the individual has. Which might not be the case with other companies.
So if we're seeing increasing number of candidates being consistently rejected at multiple places the question "why" is a valid one.
In other words, the root issue is not "surveillance pricing" but lack of competition.
No worrying where the next meal will come from, if there's going to be enough crops for the next few months, or if you'll be able to find an animal to kill large enough to feed you but but not large enough to kill you, if you can protect yourself against predators, or aggressive neighboring tribes, if you will be able to find/maintain a shelter good enough to protect you from the elements, esp in extreme cold or hot climates. If you'll be able to make enough shoes to earn enough to sustain yourself and the family, while competing with other shoemakers for a limited demand and limited materials, and million other things.
> In fact, quantitative studies revealed that the average adult hunter-gatherer spent about 20 hours a week at hunting and gathering, and a few hours more at other subsistence-related tasks such as making tools and preparing meals (for references, see Gray, 2009). Some of the rest of their waking time was spent resting, but most of it was spent at playful, enjoyable activities, such as making music, creating art, dancing, playing games, telling stories, chatting and joking with friends, and visiting friends and relatives in neighboring bands.
I'm surprised the author didn't add that they also didn't suffer from obesity or dental cavities or cancer (which is mostly because living past 30 wasn't invented until like 14th century).