But I much prefer this approach over allowing models to develop their own hyper-optimized information exchange protocols that are are black box to humans, and I hope things stay this way forever.
294 karma · joined March 24, 2016
But I much prefer this approach over allowing models to develop their own hyper-optimized information exchange protocols that are are black box to humans, and I hope things stay this way forever.
As for device telemetry, my experience has been that most companies don't rely too much on it. Any heuristic used to identify bots is likely to have a high false positive rate and include many legitimate users, who then complain about it. Captchas are much more common and effective, though if you've seen some of the newer puzzles that vendors like Arkose Labs offers, it's a tossup whether the median human intelligence can even solve it.
If anything, it's unethical for companies to dictate how their customers can access services they've already paid for. If I'm paying hundreds of thousands per year for software, shouldn't I be allowed to build automations over it? Instead, many enterprise products go to great lengths to restrict this kind of usage.
I led the team that dealt with DDoS and other network level attacks at Robinhood so I know how harmful they are. But I also got to see many developers using our services in creative ways that could have been a whole new product (example: https://github.com/sanko/Robinhood).
Instead we had to go after these people and shut them down because it wasn't aligned with the company's long term risk profile. It sucked.
That's why we're focused on authenticated agents for B2B use cases, not the kind of malicious bots you might be thinking of.
This actually creates an evergreen problem that companies need to overcome, and our paid version will probably involve helping companies overcome these barriers.
Also I should clarify that we're explicitly not trying to build a playwright abstraction - we're trying to remain as unopinionated as possible about how developers code the bot, and just help with the network-level infrastructure they'll need to make it reliable and make it scale.
It's good feedback for us, we'll make that point more clear!
If you're trying to build an agent for a long-running job like that, you run into different problems: - Failures are magnified as a workflow has multiple upstream dependencies and most scraping jobs don't. - You have to account for different auth schemes (Oauth, password, magic link, etc) - You have to implement token refresh logic for when sessions expire, unless you want to manually login several times per day
We don't have most of these features yet, but it's where we plan to focus.
And finally, we've licensed Finic under Apache 2.0 whereas Browserless is only available under a commercial license.
Thanks for the feedback! I just updated the repo to make it more clear that it's Playwright based. Once my cofounder wakes up I'll see if he can re-record the video as well.
I don't think the dead internet theory is true today, but I think it will be true soon. IMO that's actually a good thing, more agents representing us online = more time spent in the real world.
MrBeast has always been clear that his goal is to make the best videos in the world. Not to be the most nurturing place to work, or the most philanthropically minded. This document makes that clear. It shouldn't come as a surprise to anyone that in becoming the best in the world at youtube, he's had to become an extremely toxic individual.
There's been a lot of "chat with your x" projects and the value prop always eludes me.
To use the example in the repo, if I want to know what image encoders are supported, I would do a repository search for the "Encoder" keyword to find where they're defined. Then I'd be able to see all the encoders that are supported. That takes me about 10 seconds - why would I want to use a chatbot to do this instead?
What assurances does perma cc give that it will continue to maintain its index for the foreseeable future as costs increase? Wayback machine is maintained by a non-profit with a charter and >$30m in annual donations/revenues. As far as I can tell Perma is maintained by a single entity (Harvard Law Library)
Not trying to be negative here but genuinely don't see how a SPOF is better than link rot.
In 6 months to a year we'll really start to see the outcomes of those employees that know how to use these tools and those that don't diverge, and companies are going to pour a lot of resources into providing training and access to them.
Upper limit depends on the model, Llama 2 is 4k including the prompt.
In terms of cost - just ran our deployed cluster through GCP's pricing calculator and it's about $300 USD per month. Definitely not cheap for individual use, but pretty affordable for enterprise use. Running the 40B parameter version will be significantly more.
In the meantime it uses GPT4all when running locally so you can technically deploy it as well, but it's not very good.
https://huggingface.co/sentence-transformers/all-MiniLM-L6-v...
You can also specify a specific embeddings model from SentenceTransformers to use in /server/.env
Are you using it with input docs or without? Locally it uses GPT4all which isn't nearly as good as Llama or Falcon. I saw a project that is docker for Llama 2 so we might use that instead!
While it would be nice to have the data set Meta used I think open sourcing the weights is good enough.
Thanks for the feedback! We’ll include a demo soon.
Maybe I'm just to used to all the marketing speak out there where "unifying" has been co-opted to mean basically nothing. e.g. "We unify technology with human potential!"
One piece of feedback: the project sounds useful once I read into it a bit more but the headline is confusing since "unifying" can mean many different things.
We also don't expect companies to customize the functionality, just to self-host it or use the cloud version, or use it for personal projects.