It seems like they should have downsampled from an actually large dataset rather than generatively upsampled from a small dataset. Unless I'm missing something?
1,341 karma · joined March 25, 2012
@GradySimon on Twitter
It seems like they should have downsampled from an actually large dataset rather than generatively upsampled from a small dataset. Unless I'm missing something?
While differences in spending surely have some effect, spending isn't the only way to get the word out. I didn't measure, but I feel pretty confident that no more than 60% of the total communication about Prop 22 that I heard from any source was pro. I think it could have been a minority actually. For every PR firm the pro side hired, there were thousands of unpaid activists on the con side (which does say something positive about the con side!).
I noticed you brought out the "lawmakers can't repeal" line. This is true, but also effectively standard for ballot measures in CA. In fact the default is that you can't repeal or change the law under any circumstances. This was one point that initially swayed me to vote "against", but once I learned that this was basically how all ballot measures work (including ones you support, I'm sure), I swung back to voting "for".
For me, this was a difficult call, but I voted "for" on the basis of the net good for the entire pool of current drivers, admittedly at the expense of the smaller number of drivers who would have been able to drive under a better deal had this passed.
The net effect this would have had on driver welfare is far from obvious, so it always bugs me when people assert that the only reason this passed by a 17 point margin is that people were tricked and confused by the ride sharing companies.
To me, neoliberalism is better characterized as balancing market and non-market approaches to addressing social problems (more market-oriented than progressives, more comfort with government intervention than conservatives).
Also, sorta weird to blame Silicon Valley attitudes for a failure of the Texas government. Yeah a couple of companies from SV are opening offices in Texas, but this problem has been at least a decade in the making and it's not like SV is a dominant force in Texas.
This kind of model could be used as the teacher in a distillation setup too though. Then faster training of the teacher is actually a huge benefit since it speeds up model development iteration cycles.
But even if it weren't practical to use in production in any sense, I'd argue there's value in doing the basic research of exploring design space of architectures in this way. This came out of a research team at Google. It may inspire and inform smaller, more practical architectures.
The advantage, as they show, is that the model can train to a given level of performance much faster with a fixed amount of computing power compared to an architecture that uses all parameters on every step. This might be because it allows you to have a very large number of parameters that can store a lot more specialized information without incurring as much of a computational cost. Of course the downside is that you end up with a very large model that literally won't fit in a lot of environments.
Also, tangential, but I think the advice to brush your teeth right after is not right. Not a dentist, but IIUC, after exposing your teeth to acid, the enamel is in a weakend state and then the abrasion from brushing will do more damage. Rinsing with water would be helpful. Brushing before can also be helpful since the fluoride temporarily protects the enamel.
It runs a machine learning model in your browser to convert the text into points in a high dimensional space, and then it projects those points down to 2/3D.
Right now you can tell it to visualize post titles or comments from any subreddit or load an hourly updating snapshot of Twitter.
You can also view your own data in it by selecting the New Nebula option. The data never leaves the browser, which also means the ML models are run in-browser (via tensorflow.js). This part might be slow and only works in Chrome unfortunately.
If you're interested in this kind of thing, I'd love to hear from you! Here or by email (grady.hsimon at gmail)
Would love to hear about others.
- Routing customer support requests - Understanding freeform user feedback - Understanding the memetisphere of social media - Automating content moderation on social platforms - Bringing order to large document archives
I think it's possible to build a code-free, interactive interface that enables all of these things (though it may be best to focus on a single vertical).
Hit me up if you're interested in any of this. grady.hsimon at gmail.
Things are looking up!
I'm an ML engineer focused on NLP applications. Contact info in my profile if you ever want to chat, e.g. about different approaches for estimating document similarity.
Re-frame in particular is a gem. It's as though someone tried the React/Redux stack, thought long and hard about actions and selectors, and realized that with one or two more pieces, everything falls into a beautiful, purely functional harmony.
You'd be surprised how easy it is to get a model that performs as well as what you see in the video. And it's even easier now that people have built great libraries for fine-tuning generative language models.
I encourage you to try it yourself! There are many interesting extensions for people to explore:
- Use bi-directional context (vanilla GPT-2 only sees backward context)
- Integrate with semantic analysis tools.
- Experiment with different context representations. You condition the model on an arbitrary sequence of N tokens. It's not necessarily the case that you should spend that whole budget on the N tokens that came immediately before. What about including the imports at the top of the file? What about the docstrings for functions that were just used? What about the filepath of the current file?
Don't look at something like this as though watching your job be automated away. Look at it as a tool that you can master and use to move up the stack.
It's part of the job to continually incorporate new capabilities and lever yourself up.
While I don't doubt people have shown that various transformer models have certain limitations, I'm pretty bullish on transformer models in general.
Here's a post exploring the application of transformers to symbolic mathematics for instance: https://medium.com/analytics-vidhya/solving-differential-equ...
I'm not so sure it's unnecessary. I (and my employer) find it very useful that I have so many colleagues within walking distance of my desk and that we have centralized meeting, event, and dining infrastructure.
I also find it very useful to have so many other employers nearby that I could work at if this job doesn't work out. My employer finds it very useful to have such a large local labor market for the kinds of roles they need to fill.
I'm not opposed to income and property taxes on these marginal increases in productivity that the tech industry gets in the Bay Area, but I strongly disagree that the industry's large presence is somehow inefficient. The tech industry actually uses the resources of the area more efficiently than other industries because of these network effects.