70 karma · joined November 19, 2022
For example, I make an open source Firefox web extension for filtering internet content with my own classifier. That literally would not be able to exist without being able to be trained on web content, much of which is copyrighted. Requiring that I somehow either a) use only attributed data or b) detect and not use copyrighted content when trying to build something representative of my source distribution (e.g. the web) sounds like a recipe for a poor outcome. Now maybe my addon isn't your cup of tea - but what if you found out that the next generation of uBlock Origin etc. could not be as effective because of legislation because it wanted to use an AI model? Legislating too heavily around this area will, I believe, have a tremendous chilling effect for small businesses and open source folks trying to innovate in AI.
I've also worked commercially in the creation of two closed source machine learning models, but the domains were restricted enough that web content was not a particularly helpful input. One did all right, and one did not. Seeing bets succeed and fail gives me appreciation for the long-term and uncertain bets that OpenAI has been making for ages finally coming to fruition. I think without businesses being willing to make those bets the GPU-hours would have been hard to pay for.
I've wondered if potentially a different way out of this is not restricting the use of copyrighted material in the training process itself, but rather to instead only consider the created final works. Of course there are thorny problems there, too, but I don't see that having the same chilling effect on research and probably a lesser effect on business as well. One thing I think is clear though: we've reached a tipping point in the US similar to 1998 when the DMCA was legislated where the technology is forcing us to think carefully about what copyright means.
So I have question for those on HN who have meaningfully worked in the creation of not just AI-generated content, but in the creation of some AI model that others use freely or commercially: what seem like promising paths forward here? Or to those working in copyright law (like @williamcotton): how do you see the status quo and potential paths forward?
I am less pessimistic about the job outlook. As both a developer and a hiring manager, I do not think the need for junior developers goes away with AI. Managers aren't magically all just writing code by whispering into the AI and having it produce their product.
A few questions/thoughts; please keep in mind I do not know the ins and outs of your situation:
- You've indicated you're sending out endless applications. As a hiring manager, one thing that particularly stands out is when the applicant shows me that they are interested in my company's specific job posting - so quality over quantity may be key. Are you including a cover letter that shows you've looked up the company a bit and are able to make a convincing argument you'd be a good fit for the job? It might seem old-school but the number of applications I see with a well-sharpened one page resume and convincing cover letter are less than 5%.
- Have you sat down and reviewed your resume with a hiring manager to help sharpen it up?
- If you are getting as far as rounds of interviews but are consistently getting rejected, it may be worth it to double check your references. Perhaps one of them feels differently about your performance than you'd expect, and is burning all of your interview hopefuls. (I encountered this on the job just this past week.)
- Have you sought and received helpful feedback from job applications where you did not succeed? Ideally this sort of feedback can help keep you from painfully repeating things.
- Are you willing to work onsite rather than remotely? Are you willing to relocate someplace less glamorous? If so, I think this may open a few new doors, and may provide opportunities to build experience.
- Regarding starting your own business. You've mentioned that you do not have a burning passion to be programming all the time. If you are considering starting your own e.g. consultancy, I'd like to note that - successful or not - one thing I've seen as a strong common factor in all entrepreneurs is that they have a strong passion or dedication to their mission and often their craft.
- Switching career paths. I do not know your situation, skills, or true expertise - but your post sounds as though you are deeply discouraged. I'd recommend seeking counsel from family or friends who know you well - they may not know software development well, but they may be able to provide perspective.
Can you tell us a bit further about your experiences and specifics on things that have worked or not?
You might get some more targeted responses here if you could give us an example of the types of roles you're applying for. As noted some fields are tougher too so I think general advice is only going to get you so far and specifics may help in so far as you are able to share.
But as a member of a small-ish company, I remember there was this time the sales VP insisted that we MUST have this product out by a certain tradeshow. Engineering worked hard and delivered, and although the thing had warts, it gave our company an advantage: the current market leader showed up at the next tradeshow a couple months later (e.g. "20%") with a similar product and price point, but our lead in time meant that we were able to exist as a strong competitor. This single product launch and subsequent sequels lifted our company for quite a long time.
So now when the sales guys insist we might need something by a certain date, I carefully evaluate rather than dismiss out of hand - I've learned to trust their role as well.
It's true that they have limitations and are not always the right tool, but it's often valuable to realize that you implicitly have a state machine already whether you wish you did or not.
It targets pornography, borderline sexualized content, as well as other graphic images that are not safe for work (NSFW) or may be inappropriate for children.
And then more details as the other commenter pointed out:
https://github.com/wingman-jr-addon/model#datasetYour response is interesting because it tells me you maybe expected it to be in a different spot - was there a specific spot you were looking at? Might help me improve the descriptions.
It's been quite a journey since then! I developed the bulk of the code but have had a bit of community help as well. There is now a small but steady user base, and I've learned so much.
* NSFW content detection is a hard task - I believe much harder than just general image recognition, and whereas many think of of image recognition as "solved", I would argue that NSFW content detection is far from it.
* Having your own dedicated toolset to assist in creating and managing the dataset is invaluable. I was surprised to see that as a single individual it was quite feasible to create a fully-curated dataset in the hundreds of thousands.
* Servers are more different than laptops/desktops than one might think. I had a good time setting up an old Supermicro with K80's on a budget.
* For this problem at least, even several hundred thousand images was still not enough to see an advantage on training from scratch vs. tuning a pre-trained model.
* Fun ways to slice up GIFs at a "transport" layer to allow for good-enough filtering on frames.
* Character encoding detection is a hard problem, and the existing filtering API's don't do a good job of helping the developer down the right path.
So with development being relatively stable for quite some time, why share now? Well, I've kept the addon fairly "primordial" in the sense that I haven't tried to cater too heavily to narrowed use cases yet. Three general use cases seem to be represented based on user feedback - there are others but so far this is what is being said:
* Adults casually enabling it for daily browsing. Think things like browsing stock photo sites, Google Images, etc.
* Adults struggling with pornography and looking for tools to help.
* Adults looking for an extra safety net when their kids browse the web.
I'm contemplating adding more specific feature sets in one or more of these areas, but thought that it might be a good time to put it out there and get some perspectives. The tech is far from perfect, but it seems that it is good enough that it is helpful for some users.
It's also my hope that there's potentially some things to share here that the HN crew might find of interest. (Although I can assure the UX is not currently one of them!)
* The addon itself (https://addons.mozilla.org/en-US/firefox/addon/wingman-jr-fi... and https://github.com/wingman-jr-addon/wingman_jr) - maybe it's not your thing, but if you're like me, there's a good chance somebody in your family might find it useful
* The model (https://github.com/wingman-jr-addon/model) - I've tried to make a competitive model for its size, such that enterprising individuals can try this as an alternative to paying for API calls. It does not use NSFW.js as a base. Both .h5 and TF.js models are provided. Maybe it'll be good enough for your use case?
* Real world examples of the webRequest.StreamFilter API. In particular, the bit about character encoding is probably worth a short read if you're thinking of using this API yourself. See https://github.com/wingman-jr-addon/wingman_jr/blob/79a1a882...
* Examples of image-based logging for Firefox.
* Some fun GIF parsing stuff!
Thanks!