Batch computing and the coming age of AI systems
hazyresearch.stanford.edu
hazyresearch.stanford.edu
Sure most of the work in the computing world is done by scripts (written by humans).
Sure if an AI can magically write/manage all those scripts that would be great.
I don't really see anything beyond someone creating an artificial divide in computing that I don't really see as relevant ("batch" vs using a literal app with an AI feature).
Of course most computing doesn't require a person to be interacting with it.
I'm being genuine here; please clarify if I'm just lacking context.
And then big cos will jump the bandwagon and we’ll have a new trend!
IntelliBatch: Intelligent Batch Processing
AutomaBatch: Automated Batch Processing with AI
I’m getting some real good vibes here. We’re getting close. I’m sure actual enterprise actors can come up with something even worse.
1. What if we used foundation models to write code?
2. But the code might not work...
I wouldn't take it as an informational article in the sense you're looking for. It's more like a press release, and pretty common in academia to release these fluff blog posts.
For anyone who doesn’t know what a PI is (I didn’t when I first got a job at a university).
Most LLM output is worthless before validation. Be it code, problem solving or question answering, they can all be wrong at any time.
But then they explained that one could use these models to generate the code that would process the financial data.
Sounds interesting, but yes the question is how do you validate this? Do humans write the test cases to ensure that, for example, ACH files are being processed accordingly? What about edge case detection? Self-correction in real-time? Many questions left unanswered.
We don’t generally build or interact with continuous control systems, with one exception: Organic things
So humans, pets, “nature” are all “streaming” systems because they don’t have a set of discrete states to move between and adjust to feedback loops. You can reduce or segment organic action into states, but this is artificial compression.
The future IMO are streaming systems (with and MDP/REPL/OODA/RL… etc controls), as batch isn’t the way intelligent systems actually behave.
The problem I am facing now is how to filter out the mislabeled training data generated with GPT3, also around 10%. Manual validation is out of the question at this scale, and it seems to be the most interesting training data if I could have it validated. I tried GPT3 fixing itself, only works partially. What I did was to train a model and use that model to rank my data, and set a threshold. That's how I got the cutoff around 10%. I can use the easy part, but what to do with the interesting boundary region data that might be mislabeled?
It makes a lot of sense to use an LLM to understand data then generate code to do mass extract, but the article is really oblique and it’s hard to tell what’s they’re doing
Batch processing is a term from HPC systems, which are typically multi-user, and have to share computing resources by means of a workload manager (eg. Slurm, PBS, UGE etc.) Workload managers function by defining queues, queues have associated resources (eg. number of CPUs, particular network hardware, accelerators, s.a. GPUs and so on). Users interact with workload managers by scheduling execution of their programs on particular queues (because they want the resources associated with those queues).
The interaction between the user and the workload manager is described as "submission of a batch job", where this is contrasted with running a program interactively (s.t. the user can simply start the program and observe its behavior right away). I believe, these are the batch jobs the article is talking about. These kinds of jobs are very common in HPC, and I'd imagine in ML performed on HPC resources.
That’s what it’s getting at- setting up routines for unattended execution in batch.
BTW Chris Ré did stuff with UW Madison /HTCondor and definitely knows what batch computing is