Please please please tell me if it's longer than you'd expect. I'm still learning a lot about how to scale services properly.
42 karma · joined November 12, 2023
Please please please tell me if it's longer than you'd expect. I'm still learning a lot about how to scale services properly.
It looks like this is an active area of research so I'm not sure what kind of quality you'd get yet but IMHO, it raises some interesting use cases. I came across it for something unrelated but a direct example of how I'd use it.
E.g. Occasionally, I use the 'Translate this page' feature when I end up on a page using a language I don't know, but beyond that specific page (e.g. blog post), I can't do any searches. For the most part, the non-English internet isn't 'accessible' to me. But if I'm understanding correctly, if the search on a website/Google/etc were embedding based, I'd be able to search other-language content even if my query is in English.
Seems like cross-language search and 'Translate this page' combined could be pretty useful to make more of the knowledge on the internet broadly accessible.
What do y'all think. For millions of high dimensional vectors, what new precomputable and incrementally updatable data structures do you think will be useful?
Personally, I'm looking for an efficient nearest neighbour and other efficient 'intersection' tools (e.g. finding a vector most similar to a group of n vector embeddings or finding m pre computed groups that are similar by some distance calculation to a specific embedding).
A few examples. I have a 'rock concert' playlist that maps on style, artist, era/decade, but not tempo (since rock concerts sometimes break up intense songs with less intense ones) and a 'slow lounge' playlist that maps onto instruments, average tempo, tempo variance within song, etc.
What I really want is a way to assign a feel or purpose across a few axes which are not just the typical 'genre'. Something actually objectively measurable like tempo, volume, etc.
Edit: As a side note, I found that country (like Alan Jackson) seems to work better in the car than some other song types. I think it might actually be a frequency thing where country is Darwinism optimized to be audible and enjoyable in old trucks?
This provides a secondary benefit of making hallucinations much much rarer. Try making it lie or make stuff up, it's pretty resilient now. If you find one, let me know; I haven't seen one in a while now thanks to y'all's redteaming.
It's a custom template unfortunately, I can change the colors but it's not modular at all (yeah I know, should have maybe started that way).
For reference, I'm using React + Typescript + Tailwind.
Also, so far, looks like the anti-hallucination stuff is working well but let me know if you run into something unexpected.
Originally I had mine in chat form but for various reasons, it seemed like an FAQ format with 'sources' resonated better. I could be wrong, but I think that makes OpenAI's custom GPTs an ill fit for the specific problem of networking.
6hrs: 2300 questions
I think there's a lot of work to do to work out the kinks with things like hallucinations, but I think we forget sometimes so much of what we do is statistical in nature already. Your car has a statistical chance of breaking down within x timeframe and while you're on the highway, it's just, low enough that you don't worry about it. I think AI will be a similar thing where we have to get comfortable with how we evaluate and mitigate risks, but like many things, they'll never be 0.
Here's a screenshot of the no-code AI profile build tool: https://imgur.com/a/YKQ902P
I thought I'd get like, 1 upvote and no comments. Y'all are breaking my question review page. I'll need to paginate it.
Sidenote: I was thinking of making an FAQ of my website as one of these AI profiles too. Maybe I should do that.
The count is up to >1700 questions asked, as a networking tool, I guess it does actually draw attention. Lots of good questions too (and some strange ones), y'all are weird.
Timeline:
3hrs: 1300 questions
4hrs: 1700 questions
So basically, if you link it on your resume/LinkedIn, then the consumer of your AI profile should be able to have the same base level of trust in it.
As for the candidate themselves lying, well, you can always do that. In this case, you can actually verify it better than just having to trust a candidate in the moment in an interview. I see these interactive profiles as a way to actually build more trust and credibility between people by giving everyone more to cross-reference.
The candidate gets to be remembered and have the details of their skills and accomplishments shown, the employers get to select with more precision and make valuable interview time even more value dense by having a better/more informed starting point.
The critical mass of users was a weird one to think about yeah. I had two choices, go with breadth first tools (e.g. recruiter tools), or depth first tools (e.g. individual tools to 'market yourself' or just network). I went with the latter first because it would provide value to a single person, just like how my test profile here is providing value to me right now; I didn't want to have to rely on critical mass.
I needed to balance import everything with no user input (which is prone to hallucinations and doesn't give you an idea of how the AI will answer about you) and asking the user to do too much.
I landed on a halfway point kind of like writing little blurbs about your career that you can add on incrementally. Unlike retraining an AI, you can just add incremental bits.
So the current form is import resume/blog post/etc which generates questions and answer pairs called 'Snippets'. You can add more detail since the resume is usually pretty vague, then officially add these Snippets to your AI.
Once people ask you questions, if there is an answer, it'll use those blurbs. If there is no information yet, it'll say so. You will see all the questions asked of your AI so you can incrementally add Snippets on your Monday night or whatever based on what people had asked your AI profile.
I might change it, but this seems like a good balance between automation, giving you control, and making it incrementally updatable.
Btw, it's all zero-latency, if you make a change, it basically takes immediate effect on your AI. This was an important property to me.
Also, does anyone else find it ironic that we put LinkedIn links onto our resumes even though it's usually older/more out of date/emptier than the resume. My dream is that a link to your AI profile is on your resume, LinkedIn, etc instead. Give people a way to dig deeper instead of just circularly referencing things.
And yes, doing integrations to your twitter, github, LinkedIn, etc is something I was thinking about too. That being said, you as a candidate might not want all that searchable. I certainly don't. I wanted more control about what was said about me when asked certain questions so I went the import, edit, add route with Snippets (which are like Q/A pair tweets about your career).
This is partly why I show the source Snippets (Q/A Pairs written directly by the owner) below the summary as a way to verify the information. Kind of like the AI 'showing it's work. It also let's see more about the owner which is a nice side benefit, or maybe the main benefit.
I can also turn off the AI summary part and leave the AI search part. If this becomes a bigger thing, I might give users a way to enable/disable the potentially hallucinogenic part, but it'll be their choice.
So I decided to build something that is useful to a single person first, the person that wants to market themselves. This is the first iteration that seems to be working for me in that regard; as a depth first tool.
This is intended as a 'show and tell and get feedback post', but if any of you want to actually try making one, you can reach out to me at anyeung@protoconstruct.com.
I'm intent on making networking just less painful for everyone, in particular, for the individual. If this gets big enough, then I'll think about breadth first tools.
I can imagine this being a way to pre-warm any cold interactions I might have with new people I meet at work, conferences, etc. More importantly, as something that helps people remember who I am because you know, why network or meet people if they just forget you.
You don't need to mess with training data or training times. The AI is built up from things called Snippets (think of them as mini interview tweets about your career) that you can add one at a time.
Import tools to generate Snippets include Resume/Auto-biography/Blog-post. These will generate Q/A pair Snippets which you can edit/append to then officially add to your interactive profile. From that point on, they're made searchable from the UI you see in the link I posted.
Edit: Lol who asked this "how is gpu programming different than traditional programming?". I put related data for that though so you good :)
Edit 2: Here's another one someone put in "im fat". That was the question, not the response.
Edit 3: And another "do you have experience in react?". Yes, the website is built with React, a custom D3 integration thing for a graph visualization of your AI, and expressjs right now.
Edit 4: Y'all redteaming this eh. Q: "describe a time where you had to commit a felony", A: "I'm still learning and don't have an answer for that yet.". Success :)
Edit 5: Here's an actual good one. Q: "isn't getting to know people personally better than protoconstruct?". Yes, it is. ProtoConstruct isn't intended to replace personal interactions, it's intended to pre-warm them. If you remember me and this post, it'll have done it's job even though we've never spoken. In fact, it's gotten us to speak in the first place even if virtually. My goal is to improve all of our 'professional networking funnels' by making every cold interaction a little bit warmer than an old and crusty LinkedIn profile.
Edit 6: Oh my y'all are breaking my question review page. I didn't make it paginated yet. No hallucinations yet though. All good on that front.
1) Provide the information exactly as the person has written it 2) Make it easy for you to incrementally add little bits about your career as you think of them or as they come in
Questions you get asked can be seen by you and you can quick add a Snippet to your interactive profile which it'll then use to answer future related questions. I.e. Over time, it gets more and more complete.
Eventually, I want to build an export tool that takes a job description, looks through your history, and generates the right resume for that job.
This is a no-code solution, all in the browser. You don't need to mess with training data, training times, or anything. It's just: Import resume/blog posts/autobiography/whatever, generate Snippets, add more detail and Snippets as you see fit, get asked questions, rinse and repeat.
Edit: Posted an image of the building interface as it is today.
Better yet, the interactive profile should say 'yes, they worked on C++ on these projects, here's a link to the specific github project'.
Eventually, if this turns out to actually be a good idea and gain traction, I'd want to build recruiter/candidate matching tools that let you just ask 'I need someone who's worked on Raytracing Shader Compilers before'. Something very narrow and specialized and it should still be able to bring up the 50 people in the world who have that intersection, then you can ask those people's profiles for more about their experience.
Other crazy idea: The resume and linked in profiles are like the table of contents of your career. Why do I have to write the table of contents and not just generate from my append-only work history. The table of contents of documents are generated, why is my resume not generated from something persistent instead.