541 karma · joined December 21, 2009
[ my public key: https://keybase.io/arturventura; my proof: https://keybase.io/arturventura/sigs/WmemHKEaIAvrvrPnKeFCKIoCUDLrhEiYcDAY11mFhDw ]
Its a cool project that gave me immense pleasure to built, however its unfortunately a intellectually masturbatory one, because although the tech is cool, I haven't found a cool application for it. If anyone is interested hit me up.
Location: Lisbon, Portugal
Remote: Yes
Willing to relocate: Maybe
Technologies: Django, Python, Vue, C, Java, JavaScript, Typescript, PyTorch, Nordic Firmware Development
Résumé/CV: https://www.surf-the-edge.com/wp-content/uploads/2024/01/Resume.pdf
Email: artur.ventura@gmail.com
Website: https://www.surf-the-edge.com/
My name is Artur Ventura, I'm a Senior Software Developer Engineer with a cross functional expertise in every level of
software development. Developing software on a professional level, in scenarios of
hundreds of thousands of users for 13 years. Focused on Applied Artificial
Intelligence, in particular Natural Language Processing pipelines. Just finishing the sale of a company and looking for a new challenge.> running on a single 8XA100 40GB node in 38 hours of training
This is a $40-80k machine. Not a diss, but I would love to see an advance that would allow anyone with a high end computer to be able to improve on this model. Before that happens this whole field is going to be owned by big corporations.
Search in the web is becoming problematic. Quality is decreasing and competition is extremely hard.
Some time ago I came to the realisation that the biggest strength of Google might also be its Achilles heel. Google is forced to create a list of links because that's the main vehicle where they drive profit from. If you were to send a question like "Who is Barack Obama?" you still will get a list of links although google knows there is a canonical answer.
However, if you were to build a new search engine from the ground up you would need to build the infrastructure to crawl the web, crawl it, index it and build the interface. That will take a lot of money and time to test one idea. And there are multiple possible attack vectors to Google's business model (privacy, subscription model, modality, etc.). You might get the chance of testing one of them, and if that fails, starting again is super expensive so you might not be able to.
My idea is to have a single open web index database, continuously updated so that you can apply ranking and embedding algorithms to it. This would reduce the cost of entry, and enable developers to build competitors to google on top of it, or create new products in the search space (for instance, a search engine for clothes). I don't know if this is interesting for anyone but if it is, hit me up.
The problem is that if you were to build a new search engine from the ground up it will take millions in infrastructure, and a lot of time for you to test one idea. And there are multiple attack vectors to Google's business model (privacy, subscription model, modality, etc.) however you might get the change of testing one of them, and if that fails, starting again is super expensive so you might not be able to get funds to do it.
My approach then became to build something that others can build on top of.
I'm currently using common crawl but my main problem is that I need to build a small toy to test it and even processing common crawl is crazy expensive. Just a single snap are 150 Tb, so this needs to be process on metal, or you're gonna pay a hefty AWS bill.
Willing to relocate: Maybe
Technologies: C, C++, Java, JavaScript, Typescript, Python, Lisp, Perl, Golang, SQL, PyTorch, TensorFlow
Resume: https://www.surf-the-edge.com/wp-content/uploads/2022/07/Res...
Email: artur.ventura@gmail.com
LinkedIn: https://www.linkedin.com/in/artur-ventura-48500314/
Github: https://github.com/nurv
I'm a Software Engineer with both backend and frontend experience. My core competence is in Machine Learning working with Natural Language Pipelines.
Remote: Yes
Willing to relocate: No
Technologies: C, C++, Java, JavaScript, Typescript, Python, Lisp, Perl, PHP, SQL, Swift, PyTorch, TensorFlow
Resume: https://www.surf-the-edge.com/wp-content/uploads/2022/07/Res...
Email: artur.ventura@gmail.com
LinkedIn: https://www.linkedin.com/in/artur-ventura-48500314/
Github: https://github.com/nurv
My professional background is in artificial intelligence. Currently I’m the Tech Lead at Bond Touch, and I’ve worked at Unbabel, a YC company, as an artificial intelligence engineer on the Machine learning team. I also have some interests in quantum computing, financial modeling and robotics.
Being able to unload and reload javascript. The initial idea was to write the website inside the website, but at the core level it requires having something akin to process isolation for javascript. It also requires the dom to be isolated.
Implementing 9p2000, and share resources across browsers. I’ve been reading about the ideas of plan 9 and i would like to implement something that allows me to connect point to point to other browsers and mount their FS into mine so we can share resources.
One of the cool results that I got was that since the dom is not directly changed (each process/worker has its own partial dom and every time that it changes it a delta is sent back to the main thread for sync) it allows javascript to be running somewhere else (another browser, back end server) and sync’ed back (much like vadaain, but more agnostic).
Most of the code was inspired by the linux kernel (which gave me a reason to go learn its internals) and is kinda nasty at some points but is written in typescript as some of you have already mentioned. Someone might find it interesting even if just for the educational purpose of
I was stuck at parents for xmas and I picked Tannenbaum “distributed systems” and “Modern operating systems”, which gave me an idea of running a "kernel" on a browser. It was more of an academic exercise than anything else, but my intention was to have a the following:
Being able to unload and reload javascript. The initial idea was to write the website inside the website, but at the core level it requires having something akin to process isolation for javascript. It also requires the dom to be isolated.
Implementing 9p2000, and share resources across browsers. I’ve been reading about the ideas of plan 9 and i would like to implement something that allows me to connect point to point to other browsers and mount their FS into mine so we can share resources.
One of the cool results that I got was that since the dom is not directly changed (each process/worker has its own partial dom and every time that it changes it a delta is sent back to the main thread for sync) it allows javascript to be running somewhere else (another browser, back end server) and sync’ed back (much like vadaain, but more agnostic).
Most of the code was inspired by the linux kernel (which gave me a reason to go learn its internals) and is kinda nasty at some points but is written in typescript as some of you have already mentioned. Someone might find it interesting even if just for the educational purpose of it
Location: Lisbon, Portugal
Remote: Yes
Technologies: Java, Python, AWS, GCP, Azure, Docker, Linux, K8s, Postgres, Redis, RabbitMQ, PyTorch, Marian, Elasticsearch, TypeScript, Vue.
Willing to relocate: No
Résumé/CV:
- https://www.linkedin.com/in/artur-ventura-48500314/
- CV (Portuguese): http://surf-the-edge.com/CV.pdf
- Résumé: http://surf-the-edge.com/Resume.pdf
Email: artur.ventura@gmail.com
---I'm a software engineer with 12 years of experience academic background in AI, in particular NLP. Deep understanding of JVM (built the first JavaScript JVM), been working more recently with Python for the past few years, particularly applied to production environment deployment of AI. Looking for interesting projects to work on.
Again, not trying to be a troll. Not American, just trying to understand.
If you are using this as a web server persistence backend, I would agree with the first, more or less accept the second and reject the third. HTTP + JSON serialisation are way slower for that kind of job.
If you are just exposing the database using only the Postgres, in that case is interesting, however, I have concerns about how more complex business logics would work with such a CRUD view.
My experience of using it is that tells me that without a undefined type keyword it becomes very boring to use. Each time you have to define a variable you have to write all type info:
Map<String,String> foo = new HashMap<>();
If java had a inferred type declaration, like Golang, Swift, Scala, etc. This would be much simpler: var foo = new HashMap<String,String>();
But from what I think there is nothing like that on Java's roadmap. In fact, I think it was proposed and rejected in the past.I truly would like to see the community come up with a open source version of it, but with such a huge size, the language would be near impossible for one single individual to implement a compatible version.