116 karma · joined March 13, 2012
Current: Head of AI Research @ Thoughtworks
Previous: Co-Founder, CEO @ Watchful.io (Acq. $TWKS) Data/AI Eng @ FB, several other startups
I guess it just feels icky that "regular expressions" has inherent meaning (i.e. can be represented entirely by a finite automaton) which has become completely diluted at this point.
That rant aside, cool paper. The idea of bridging formal language theory with modern computational tooling feels timely. I think I would've liked to see more exploration of oracle-based costs, for instance:
* What happens when oracle outputs are inconsistent/uncertain?
* What happens as oracle interactions become more computationally expensive?
In a nutshell: Conceptual uncertainty is when the model isn't sure what to say, and Structural uncertainty is when the model isn't sure how to say it.
You can play around with this yourself in the demo!
Re: our definitions of average/long/short prompts -- we weren't really rigorous with those definitions. In general, we considered anything under 100 tokens "short", 100-300 average, and 300+ large.
Our intuition here is that the relationship between performance of the estimation and the prompt structure is less about length, and more about "ambiguity". Again, we don't really have a rigorous definition of that yet, but it's something we are working on. If you take a look at the prompts in the analysis notebook you might get a sense of what I mean: prompts 1-3 are pretty straight forward and mechanical. Prompts 4 & 5 are a bit more open to interpretation. We see performance of the estimation degrade as prompts become more and more open to interpretation.
1. The perturbation method could be improved to more directly capture long-range dependency information across tokens
2. The scoring method could _definitely_ be improved to capture more nuance across perturbations.
I think what we've found is that there does seem to be a relationship between the embedding space and attributions of LLMs, so the next step would be to figure out how to capture more nuance out of that relationship. This sort of side-steps the question you asked, because honestly we'd need to test a lot more to figure out the specific cases where an approach like this falls short.
Anecdotally - we've seen the greatest deviation between the estimation & integrated gradients as prompt "ambiguity" increases. We're thinking about ways to quantify & measure that ambiguity but that's its own can of worms.
Obviously not a 'secure' system by any stretch of the imagination but it's an order of magnitude better than storing in plaintext.
If you've been hit with an OS compromise you're pretty much SOL, but it shouldn't be so easy to grab highly sensitive data from accidentally exposed profiles.
Took much inspiration from dumpmon, but distributed it so users can choose their own sensitivity settings.
Glad to see dumpmon is still going strong :)
I think a 'real' analysis of the data at a much larger scale (probably state-wide would be a good test) would yield some interesting data on the effectiveness of this implementation in specific.
As mentioned in the post - the circle packing implementation itself is general enough to be used in other applications. The point was to show off a cool potential application for a solution to a non-intuitive problem.
As mentioned in the post - deduping is cheap and easy so introducing the third basis vector u3 could potentially be a great move away from the recursive logic that currently plagues the project.
Some things that I think would make this an awesome app:
1) IFTTT integration would be sweet. "If i get an e-mail about something, remind me to do it by putting it in my contextual note for an application"
2) iCloud integration would be sweeter. Unified notes for applications that exist in "desktop form" and "mobile form".
3) A "view all notes" option to be able to browse sans-context
4) The ability to clip media into a note easily
I have a half-baked contextual analysis implementation which I could probably spin into a high-volume twitter analysis tool. Was doing NLP analysis on unstructured data (like news articles) and extracting topics + extrapolating commonalities between sets. Could be used to pick up topics from tweets and determine if two unrelated tweets are actually talking about the same thing (without necessarily replicating the same syntax).
The results of the more-filtered-list-tool would be quite interesting, though, as you'd essentially be modeling a set of "ideal leads" and determining how close/far a set of tweets are to those models. Just figuring out an "ideal lead" model for the segments you're targeting would be an interesting intellectual pursuit.
I think I might end up building this...
You also miss out on users asking for suggestions who aren't currently using a competitor product (which IMO is a more valuable segment).
A more interesting implementation is one that takes context into account, but that would require some homemade ML work and likely outside of the scope of quick & hacky solutions.
Some notes on why not:
+ Your web-sume looks rough. As pointed out by others, there are a number of typos (i.e: "and provide an opporunity") not to mention the design itself could use work. If you are GREAT at web design/UX you should spruce it up. Otherwise, kill it and move to a traditional resume. Knowing HTML5/CSS3 today is pretty meaningless, so showcasing that is sort of pointless.
+ There are tons of issues with your resume itself (i.e: "Excellent verbal and written communication skills." despite multiple typos and unclear flow) which need to be addressed. Cut the fluff, point to recent projects & address why they are cool/why anyone should care. Anything that you did 10+ years ago that isn't directly applicable to what you want to do in the near future has no place on the resume.
+ Your bitbucket projects are lackluster. You don't follow good git branching habits, your commits are non-atomic, your code is cumbersome and unfinished in many places. You also seem to use .py files as notes in non-standard ways, introducing weird artifacts and conventions to your projects.
Some notes on how to improve:
+ Learn how to use git productively in a team environment (this means no more working directly out of master). This is a good resource to that end: http://nvie.com/posts/a-successful-git-branching-model/
+ Learn better coding habits in whatever language(s) you are most comfortable with. Your bitbucket only has python code, so learn how to do things in more 'pythonic' ways. (i.e: Don't just stub notes inside .py files. Throw them inside a README.md or keep them in a secondary utility so you don't clutter the repo).
+ Sort of back to point #1 but deserves its own category: Learn how to use .gitignore. You have tons of artifacts in your repos that do not need to be/should not be there.
If you address all of the above, you'll be in a much better position to start qualifying for entry level dev openings.
I was hiring for a telehealth startup based in NYC and got some referrals to some recent General Assembly grads. I bit and went ahead and scheduled some interviews. It was a total joke, honestly. The graduates glossed over the entry-level interview questions with a lot of handwaving (I would ask them things like "How would you do x given y?") and quoted rates upwards of $100/hr even though their total experience was the 10 week course @ GA campus.
I was so turned off by that experience that I just never even considered hiring from a 'hacker bootcamp' again. I'll echo what others have been saying as well: You can hire them, and maybe they'll perform for a while - but the amount of time you'll need to spend to get them up to speed on CS basics will more than likely not be worth the investment. You're better off hiring a recent college grad whose only experience is working with Java - at least they have the fundamentals and can build on top of them instead of backtracking.
source: a number of non-trivial node deploys
tl;dr - It's a wave that changes slightly when you look at it. Pretty cool, when you realize that every stage of the simulation is unique, temporary, and will never be re-rendered again :)
http://pasteyebeta.herokuapp.com/ http://github.com/shayanjm/pasteye
Inb4 not reading before posting