29 karma · joined February 23, 2020
I would suspect this helps make moderation models better at estimating confidence levels of ai generated content that isn’t labeled as such (ie for deception).
Surprised we aren’t seeing more of this in labeling datasets for this new world (outside of captchas)
I think what people really want is business rules and data cleaning and schema discovery.
If you had to use English against multiple source systems to and tons of joins, the sentence would be paragraphs.
Where I think there’s value is in using something like a data catalog to label business rules against a data warehouse, tied to dashboard queries and other common ones.
But that’s a hard problem and a unique model to every customer. And always changing.
I mean even having it document a best draft of what the extension code is doing would be awesome.
Unless it’s made into an extension and then you have a recursive hell.
And you don't need that pesky transformation part.
Except you really do, when you get beyond having a source system or two.
Seems like Microsoft was frustrated with the pace of movement in this space and the shitty results of agents (which admittedly kept my interest turned away from agents for the last few months). I'm interested again because it makes practical sense, and from looking at the example notebooks, seems fairly easy to integrate into existing applications.
Maybe this is the 'low code' approach that might actually work, and bridge together engineering and non-engineering resources.
This example was what caught my eye: https://github.com/microsoft/FLAML/blob/main/notebook/autoge...
the more folks experiment with this stuff, i think they'll see where it all comes together, and where some of the crossover is. but given how quickly everything changes in this space, i'm glad there's a clear focus from each team on their core strengths rather than throwing the kitchen sink of new papers at it.
that said, while there's some clear crossover between the two, i find myself using langchain for things like huggingface embeddings for local models, and other helpers that work well with llamaindex.
somewhat akin to a data warehouse and all the techniques and abstractions that go into modeling it for non-technical end users, llamaindex makes a lot of that much easier to work with as a developer. structured and unstructured data can be indexed side by side, and the auto retriever functions they've recently built out work really well once you've got data indexed in a sensible way. our next step is to put a simple UI on top of it all with filters (like a dashboard) that pass metadata filters to the llamaindex autoretriever.
these patterns may not be exactly right today, but I don't see any others focusing on this area. just throwing all of your docs haphazardly into an index and calling it a day is no different than tossing all your data into a single database schema without any rhyme or reason, and hoping your dashboards can do 'magic' on top of it.
Also Explainpaper and Elicit for research papers - hard to go back to the 'old' days after using all 3 of these.
the process is grueling and by the end you just want it over to return back to normal life. I agree fully that getting everything done before the LOI is the way to go.
Learned a ton in this process and hope to be able to use that experience to help others as they plan to go through it.
Sell project work that's strategic (ie: solving a business problem and coming with a solution), as opposed to 'We need a frontend developer for 12 months'. Better rates, opportunities to work with business instead of IT, etc.
Like any good long-term investment, you have to keep nurturing it. Continuously stay in touch with people you meet.
If you do it right, the leads come to you through natural conversations, and you don't even feel like you're selling. Without the relationships already in place, everything feels like selling, and it's exhausting and significantly harder to land strategic project work. There's no magic here, and no book or course that will tell you the secret (it's mostly snakeoil) - it just takes time and effort. It gets easier when some of these relationships lead to real opportunities, you show well, and can use that for referrals into other places as the people you meet move around.
Another route for early starters is to subcontract through established consulting firms. Our company is more of an FTE model, but when we need specific / niche skills that we don't always have in house, we consider subcontracting options. If that's something you're interested in (we're US based), I'd be happy to chat (email in profile) and see if there's anything coming up.
(edit: the rationale behind this tends to be that you can avoid the heavy lifting of ETL/transformation logic by just using a data lake - obviously not the case, as most of us know)
I've found data lakes complement DW's (in databases) well. Keep the raw data in the lake and query as needed for discovery, and load it into structured tables as the business needs arise.
Data lakes alone are doomed to be failures.