4,312 karma · joined July 2, 2010
meet.hn/city/40.7596198,-111.886797/Salt-Lake-City
Socials:
- bsky.app/profile/joferkington
- github.com/joferkington
---
(And yes, a lot of science is software. Analysis is software.)
The issue isn't a lack of economic interest.
It might be a lack of training data in addition to inherent complexity, but it's certainly not a lack of economic interest.
But yeah, if you want to feed it math and get code, it's reasonably okay with that. All LLMs I've used seem bad at understanding things that don't look like broad human knowledge. I've seen this same general issue across many different models. (And to be fair, geology, geophysics, and remote sensing are what I'm testing, and their semi-rare niches.)
It's also quite dangerous because it's not obvious that what it's doing is complete hallucinations unless you actually are a domain expert. Things _sound_ reasonable. E.g. "this is likely feature X" which _does_ exist, but is absolutely _not_ relevant to the problem or present in the input dataset.
But my current employer is pushing this exact thing (human language + scientific data + LLM -> advanced analysis of scientific data by LLM -> business decisions) and it _really_ worries me. It often gives the rough equivalent of "Start the procedure by severing the patient's aorta. Once they stop moving, you can deal with the hangnail". Just in very reasonable sounding language. And a lot of people don't know any better, because most users aren't domain experts.
I've done a fair bit in the field, but a huge part of my career has been mining old datasets and reinterpreting things in light of new data/etc.
What the article is describing isn't new in any way. But it also doesn't remove the need for fieldwork or the need for the experience of having done fieldwork to use existing datasets. Observational sciences (e.g. geology, biology, etc) where you can't easily replicate the environment you are studying in the lab are always going to hinge on some sort of fieldwork.
Finding creative ways to use existing data doesn't change that.
Sure, the pieces average 6 faces when materials are relatively homogenous and iostropic (i.e. no preferential direction to break in and no free surface nearby). However, as they note in the article, this isn't always the case. Things like mud flats and other cases with very anisotropic materials and/or free surfaces nearby don't fracture with the same average.
This is a good example of a potential metric that could be used to give some clues about overall material behavior even if all you have are the broken remains.
Fractal dimension is also pretty esoteric. However, it's somewhat widely used in geoscience, even though what we're measuring isn't _actually_ fractal. It's still a very useful comparative metric, though, because it lets us measure how complex an interface or surface is quantitatively and scale-independent.
With that said, the simple truth of it is that we know next to nothing about these ecosystems and really can't accurately estimate impacts. They're quite possibly significant, but we just don't have much info to go off of and studies like this are sorely needed.
The article makes this point, but it's relatively far in and I felt it was worth making again.
With that said, my employer now appears to be in this business, so I guess if there's money there, we can build the satellites. (Note: opinions my own) I just don't see how it makes sense from a practical technical perspective.
Space is a much harder place to run datacenters.
Typing in python within the scientific world isn't ever used to check types. It's _strictly_ only documentation.
Yes, MyPy and whatnot exist, but not meaningfully. You literally can't use them for anything in this domain (they wont' run any of the code in question).
Types (in this subset of python) are 100% about documentation, 0% about enforcement.
We're setting up a _documentation_ system that can't express the core things it needs to. That worries me. Setting up type _checks_ is a completely different thing and not at all the goal.
"array-like" has real meaning in the python world and lots of things operate in that world. A very common need in libraries is indicating that things expect something that's either a numpy array or a subclass of one or something that's _convertible_ into a numpy array. That last part is key. E.g. nested lists. Or something with the __array__ interface.
In addition to dimensionality that part doesn't translate well.
And regardless, if the type representation is not standardized across multiple libraries (i.e. in core numpy), there's little value to it.
There's not a good way to say "Expects a 3D array-like" (i.e. something convertible into an array with at least 3 dimensions). Similarly, things like "At least 2 dimensional" or similar just aren't expressible in the type system and potentially could be. You wind up relying on docstrings. Personally, I think typing in docstrings is great. At least for me, IDE (vim) hinting/autocompletion/etc all work already with standard docstrings and strictly typed interpreters are a completely moot point for most scientific computing. What happens in practice is that you have the real info in the docstring and a type "stub" for typing. However, at the point that all of the relevant information about the expected type is going to have to be the docstring, is the additional typing really adding anything?
In short, I'd love to see the ability to indicate expected dimensionality or dimensionality of operation in typing of numpy arrays.
But with that said, I worry that typing for these use cases adds relatively little functionality at the significant expense of readability.
I wish I'd done this more.
In some cases there was no way to. For example, we once woke up to find that the European half of our team had been laid off as part of huge cuts that weren't announced and even our manager had no idea were coming. There's no good way to do layoffs, but I think that "sudden shock" approach is worst of all, personally. You don't get to say goodbye in any way and people don't get to plan for contingencies at all. (The other extreme of knowing it's coming for a year and applying for your own job and then having 2 months to sit around after you didn't get it also sucks, and I've done that as well. You can at least make plans in that case, though.)
On the other hand, in a _lot_ of other cases, you do have a chance to say goodbye. Take it. This is really excellent advice. It's worth saying something, at very least to the people you really did enjoy working with.
There's a decent chance you work with some of those folks in the future, and even if you don't, it really does mean something to be a kind human.
With that said, it's still a great tool for the job because the different stakeholders can inspect it.
Pyrolysis is a less energy intensive way to produce hydrogen, and does deserve more attention. But it still requires methane as a feedstock.
Hydrolysis let's use use hydrogen as essentially a fixed loss battery. It's perfectly complimentary to seasonally variable renewables like wind and solar. Batteries have too high of a loss though time for seasonal or multi-year storage. If you can store it (big if... Not everywhere has a salt dome like Delta, UT), hydrogen really is a great solution.
Hyperspectral in the SWIR range is what you really want for this, but that's a whole different ball game.
The "smell test" takes longer than you think and often involves an actual interview.
But when the job description contains a lot of very general terms (e.g. "scientific computing") and every part of your job history is just parroting a specific term used in the job description with no details it doesn't pass the smell test.
I absolutely respect keyword-heavy job/project descriptions. You kind of have to do it to make it through filtering by most recruiters. But real descriptions are coherent and don't just parrot back terms in ways that makes it clear you don't understand what the are. You find a way to make a coherent keyword soup that still actually describes what you did. That's great! But it's really obvious folks are misrepresenting things when a resume uses all the terms in the job description in ways that don't make sense.
I kinda think we've reach this weird warfare stage of folks submitting uniquely LLM-generated resumes for each position to combat the aggressive LLM-based filtering that recruiting is starting to use. I assume people think they can do well in an interview if they can just get past the automated filtering. I'm sure some are trying to do 3 and 4 remote jobs at once with little real responsibilities, too, but I find it hard to believe that's the majority. I may be very wrong there, though...
We had 1200 applications for an extremely niche role. A huge amount were clearly faked resumes that far too closely matched the job description to be realistic. Another huge portion were just unqualified.
The irony is that there actually _are_ a ton of exceptionally qualified candidates right now due to the various layoffs at government labs. We actually _do_ want folks with an academic research background. I am quite certain that the applicant pool contained a lot of those folks and others that we really wanted to interview.
However, in practice, we couldn't find folks we didn't already know because various keyword-focused searches and AI filtering tend to filter out the most qualified candidates. We got a ton of spam applications, so we couldn't manually filter. The filtering HR does doesn't help. All of the various attempts to meaningfully review the full candidate pool in the time we had just failed. (Edit: "Just failed" is a bit unfair. There was a lot of effort put in and some good folks found that way, but certainly not every resume was actually reviewed.)
What finally happened is that we mostly interviewed the candidates we knew about through other channels. E.g. folks who had applied before and e-mailed one of us they were applying again. Former co-workers from other companies. Folks we knew through professional networks. That was a great pool of applicants, but I am certain we missed a ton of exceptional folks whose applications no actual person even saw.
The process is so broken right now that we're 100% back to nepotism. If you don't already know someone working at the company, your resume will probably never be seen.
I really feel hiring is in a much worse state than it was about 5 years ago. I don't know how to fix it. We're just back to what it was 20+ years ago. It's 100% who you know.
By "we" I mean my company and my product does not do that. That part holds. (or, well, more precisely, that's a different product that I don't work on and isn't marketed as "imagery")
But yes, some other mosaic products are specifically requested with planes photoshopped out of airports and all waterbodies a consistent artificial color so that sunglint can be automatically simulated in flight sims for training pilots.
Because that data is often a high quality dataset available for purchase, sometimes google/etc reuses those datasets.