HNHacker News
TopNewBestAskShowJobs

codelion

3,064 karma · joined February 23, 2008

submissionscomments
codelion··on The semantic web is now widely adopted
I started this thread on the w3c list almost 20 years ago - https://lists.w3.org/Archives/Public/semantic-web/2005Dec/00...

Unfortunately, it is unlikely we will ever get something like a Semantic web. It seemed like a good idea in the beginning of 2000s but now there is honestly no need for it as it is quite cheap and easy to attach meaning to text due to the progress in LLMs and NLP.

codelion··on Why we picked AGPL
What's wrong in having a CLA? It allows you to re-license your project in future.
codelion··on Show HN: Airstrip – An open-source platform to build and use internal AI apps
Is there a hosted version one can try without having to set it up locally?
codelion··on Why we picked AGPL
This resonated with us as well, we have also chosen AGPL as license for our open source project - https://github.com/patched-codes/patchwork
codelion··on Oscar, an open-source contributor agent architecture
We also tried to implement a GH Issue to PR workflow in patchwork - https://github.com/patched-codes/patchwork

It is a bit hard to get it to work reliably except for small changes.

codelion··on Oscar, an open-source contributor agent architecture
Yes, we also use treesitter.
codelion··on Oscar, an open-source contributor agent architecture
We have a few large open source projects already using patchwork to help manage issues and PRs - https://github.com/patched-codes/patchwork
codelion··on Oscar, an open-source contributor agent architecture
We have a patchflow that does that in patchwork - https://github.com/patched-codes/patchwork it is called ResolveIssue - https://github.com/patched-codes/patchwork/tree/main/patchwo...
codelion··on Back to our roots
Yeah I was hoping there will be some bit about how LLMs make many of the tasks that people used SpaCy for a lot easier and cheaper. That is a bigger threat to the project going forward I don’t see how it can exist in the same form as today.
codelion··on Why we no longer use LangChain for building our AI agents
Many such cases. It is very hard to balance composition and abstraction in such frameworks and libraries. And LLMs being so new it has taken several iterations to get the right patterns and architecture while building LLM based apps. With patchwork (https://github.com/patched-codes/patchwork) an open-source framework for automating development workflows we try hard to avoid it by not abstracting unless we see some client usage. As a result you do see some workflows appear longer with many steps but it makes it easier to compose them.
codelion··on Learnings from fine-tuning LLM on my Telegram messages
GPT-2 is surprisingly good at fine-tuning such conversations even now. I gave a talk recently on "Sparks of Digital Immortality" that covers a bit about how we did it - https://www.youtube.com/watch?v=F9-Qk86QyMM
codelion··on Learnings from fine-tuning LLM on my Telegram messages
We do this at https://meraGPT.com
codelion··on RoboCat – A Self-Improving Robotic Agent
You can do that today with https://meraGPT.com
codelion··on GitHub Copilot Chat Leaked Prompt
You need to buy the hardware (small edge device based on Nvidia Jetson) to train and run the models locally. The demos on the site are just examples trained on my own personal data.
codelion··on Bot or Human? Detecting ChatGPT Imposters with a Single Question
The particular example you shared has been prompted with chain of thought (or may be you are using GPT-4?.

This is what happens if you try directly.

You: Please count the number of t in eeooeotetto.

ChatGPT: There are 5 t's in "eeooeotetto".

codelion··on GitHub Copilot Chat Leaked Prompt
I had similar issues when training personal models for https://meraGPT.com A meraGPT model is supposed to represent your personality so when you chat with it you need to do it as if someone else is talking to you. We train it based on the audio transcript of your daily conversations.

The short answer to how abilities like in-context learning and chain—of-thought prompting emerge is that we don’t really know. But for instruction-tuned models you can see that the dataset usually has a fixed set of tasks and the initial prompt of “You are so and so” helps model align it to follow instructions. I believe the datasets are this way because they were written by humans to help others answer instructions in this manner.

Others have also pointed out how RLHF may also be the reason why most prompts look like this.

codelion··on Vulnerability Management for Go
Good to see Govulncheck doing a vulnerable methods analysis for surfacing only the relevant issues. Many app sec vendors do it now for languages like Java and .NET. I originally created the vulnerable methods analysis back in 2015 - https://www.veracode.com/blog/managing-appsec/vulnerable-met... the same idea has been now implemented by WhiteSource (Mend), Snyk etc.
codelion··on Ask HN: How do you work with Dependabot?
To address this issue we designed a static analysis that can check if the upgrade is likely to break the application. Here are some details of the work- “Effective Static Checking of Library Updates” https://dl.acm.org/doi/abs/10.1145/3236024.3275535

When using the analysis a PR for upgrading the dependencies would look like this - https://github.com/tmroberts56/java-maven/pull/3

codelion··on Hardening attack surfaces with formally proven binary format parsers
There is already a formally verified kernel - https://sel4.systems/home.pml
codelion··on Modern Portfolio Theory: A Case Study on Turnips
Good article, reminded me of another article about MPT which used it on cryptocurrencies - https://medium.com/@asankhaya/build-a-portfolio-of-cryptocur...
codelion··on Ask HN: Best books under 200 pages for developers?
Test Lean and Ship Healthy - https://github.com/srcclr/test-lean

A short handbook on developing high quality software in the DevOps world.

codelion··on Grappling with Infinity in Constraint Solvers
It is also possible to build a model of Presburger arithmetic extended with infinity. We did that a few years back, you can try it out online here - http://loris-5.d2.comp.nus.edu.sg/SLPAInf/
codelion··on Operation Rosehub – patching thousands of open-source projects
Ah yeah I had the old link, thanks for fixing. Actually we privately disclosure the problem to the developers and get it fixed following responsible disclosure and not post PRs directly.
codelion··on Operation Rosehub – patching thousands of open-source projects
we have been doing thus for a while now : https://www.sourceclear.com/blog/millions-of-program-builds-...
codelion··on TruffleHog – Searches through Git repositories for high entropy strings
Interesting project, you can also search for all kinds of suspicious commits using commit watcher: https://github.com/srcclr/commit-watcher
codelion··on Bokeh – a Python interactive visualization library
Great library, specially useful for doing streaming visualizations. I gave a talk last year at PyData Singapore about it, slides are available here - http://asankhaya.github.io/professional.html#PyDataSG
codelion··on Floyd–Hoare logic
If you are interested in this sort of thing you can also have a look at HIP/SLEEK Verification System - http://loris-7.ddns.comp.nus.edu.sg/~project/TeachHIP/

It uses separation logic which is an extension of Hoare logic for programs that manipulate the heap.

codelion··on Holding Your Breath in India
Air pollution is a big problem in New Delhi. I wrote about it almost 2 years ago when looking at the data and comparing it with Beijing - http://asankhaya.blogspot.sg/2013/06/air-pollution-in-new-de...
codelion··on QA = Time and Money. How much should you invest?
I am not sure you want to delay penetration testing until you are post IPO.
codelion··on Using Haskell's QuickCheck for Python
Neat trick using the FFI for this. Also, quickCheck like libraries are available in other languages e.g JUnit-QuickCheck for Java - https://blog.sourceclear.com/property-based-testing-for-java...
← PreviousPage 4 of 6Next →