HNHacker News
TopNewBestAskShowJobs

abraxaz

248 karma · joined July 10, 2019

submissionscomments
abraxaz··on Show HN: Every great read I've come across, compiled into a knowledge graph
Nice graph, have you considered making it a directed graph, and also assigning more explicit semantic meaning to the edges?

So for example, using turtle syntax [1], instead of

<https://engineering.zalando.com/posts/2022/04/functional-tes...> <http://example.com/graph-edge> <https://www.testcontainers.org/>

have

<https://engineering.zalando.com/posts/2022/04/functional-tes...> <http://purl.org/dc/terms/subject> <https://www.testcontainers.org/>

The semantics of http://purl.org/dc/terms/subject is given at the url itself, but in brief:

> A topic of the resource.

> Recommended practice is to refer to the subject with a URI. If this is not possible or feasible, a literal value that identifies the subject may be provided. Both should preferably refer to a subject in a controlled vocabulary.

This would be similar to how wikidata expresses knowledge [2]:

<http://www.wikidata.org/entity/Q28315661> <http://www.wikidata.org/prop/direct/P921> <http://www.wikidata.org/entity/Q750997>

Or in English:

"Go To Statement Considered Harmful"(Q28315661)'s "main subject"(P921) is "goto"(Q750997)

This also makes it easier to query [4], for example, you could get all articles covering a "goto" with the following SPARQL[5] query:

SELECT ?item WHERE { ?item <http://www.wikidata.org/prop/direct/P921> <http://www.wikidata.org/entity/Q750997> }

May help to read the RDF primer [3] also.

[1]: https://www.w3.org/TR/turtle/

[2]: https://www.wikidata.org/wiki/Q28315661

[3]: https://www.w3.org/TR/rdf11-primer/

[4]: https://w.wiki/5RW2

[5]: https://docs.stardog.com/tutorials/learn-sparql

abraxaz··on Benthos: Fancy stream processing made operationally mundane
This is an incredibly useful swiss-army-knife like tool. I have similarly found rclone to be quite useful. I'm wondering if people know of more tools of similar nature than rclone and/or benthos.
abraxaz··on Making the collective knowledge of chemistry open and machine actionable
> it's pretty difficult to build One ELN to Rule Them All given how flexible many kinds of biological experimental designs are - especially when you're working on the bleeding edge.

RDF is quite flexible and using a combination of domain specific ontologies like cheminf[1] and other top level ontologies like BFO[2] should allow you to capture most of the semantics.

[1]: https://www.ebi.ac.uk/ols/ontologies/cheminf [2]: https://en.wikipedia.org/wiki/Basic_Formal_Ontology?wprov=sf...

abraxaz··on Super-Structured Data: Rethinking the Schema
A place to start looking may be the OWL primer (https://www.w3.org/TR/owl2-primer/) and the RDF primer (https://www.w3.org/TR/rdf11-primer/)

Other resources: https://github.com/semantalytics/awesome-semantic-web

abraxaz··on Super-Structured Data: Rethinking the Schema
> Ok, fine. But I'm not sure how this helps if you have six different systems with six different definitions of a customer, and more importantly, different relationships between customers and other objects like orders or transactions or locations or communications.

If you have this problem, consider giving RDF a look - you can fairly easily use RDF based technologies to map the data in these systems onto a common model, some examples of tools that may be useful here is https://www.w3.org/TR/r2rml/ and https://github.com/ontop/ontop - you can also use JSON-LD to convert most JSON data to RDF. For more info ask in https://gitter.im/linkeddata/chat

abraxaz··on Super-Structured Data: Rethinking the Schema
To pile on a bit here, JSON-LD is based on RDF, which is an abstract syntax for data as semantic triples (i.e. RDF statements), there is also RDF* which is in development which extends this basic data model to make statements about statements.

RDF has concrete syntaxes, one of them being JSON-LD, and it can be used to model relational databases fairly well with R2RML (https://www.w3.org/TR/r2rml/) which essentially turns relation databases into a concrete syntax for RDF.

schema.org is also based on RDF, and is essentially an ontology (one of many) that can be used for RDF and non RDF data, but mainly because almost all data can be represented as RDF - so non RDF data is just data that does not have a formal mapping to RDF yet.

Ontologies is a concept used frequently in RDF but rarely outside of it, it is quite important for federated or distributed knowledge, or descriptions of entities. It focuses heavily on modelling properties instead of modelling objects, and then whenever a property occurs that property can be understood within the context of an ontology.

An example is the age of a person (https://schema.org/birthDate)

When I get a semantic triple:

<example:JohnSmith> <https://schema.org/birthDate> "2000-01-01"^^<https://schema.org/Date>

This tells me that the entity identified by the IRI <example:JohnSmith> is a person - and their birth date is 2000-01-01. I however don't expect that i will get all other descriptions of this person at the same time, I won't necessarily get their <https://schema.org/nationality> for example, even though this is a property of a <https://schema.org/Person> defined by schema.org

I can also combine https://schema.org/ based descriptions with other descriptions, and these descriptions can be merged from multiple sources and then queried together using SPARQL.

abraxaz··on Take action to protect Open Source today
> If you want to 'protect' FOSS projects you care about, take some time to find out what help is useful to the maintainers and contribute towards items that make sense to you. Joining OSI won't help those struggling projects you gain from using.

Indeed, it is an incredibly rewarding experience. Take one thing you use and like, go to it's issue backlog and start fixing/improving things - if there is nothing take the next thing, there are likely 10s of thins you rely on every day that need contributors and contributions. The first issue will be hard, the next one easier, you will be a happier person for doing it, you will make a bigger impact than starting another project you won't finish and that nobody will use, and you will become a better engineer.

Another option is to fund actual open source projects, like go sponsor python: https://github.com/sponsors/python

abraxaz··on NFT's aren't the answer to the problems of digital art
> NFTs do solve a problem, and the problem that they solve is that creatives are getting paid for their creative output.

This has been happening for 1000s of years already before NFTs.

> Never in history has anyone been able to buy shares of highly coveted art... maybe now you can?

Yes, they have: https://www.masterworks.io/

abraxaz··on Ask HN: Best Alternative to Homebrew in 2021?
Conan can be made to do what homebrew does with minimal effort, I have written some convenience wrappers around it which makes it slightly easier to use for this use case, you can have a look here: https://gitlab.com/aucampia/proj/xonan
abraxaz··on Wikidata
Last I checked mix-n-match was using CSV, while this is okay, it still would be nicer to have direct RDF ingestion. And yes, I realize the reason why Wikidata does not have it, but it is not impossible to provide, just really difficult. I would work on it if I had more time and would likely sometime in the future.
abraxaz··on Wikidata
> I’ve been considering using it for some projects, but the main thing that’s keeping me away is the concern that some moderator will decide that my data doesn’t fit and remove it.

To me the greatest value of Wikidata was making me aware of RDF and SPARQL.

In most cases, if you are relying on data business needs, it would be best to maintain your own RDF dataset and host it either just on HTTP, or on something like https://dydra.com/.

WikiData deseperately needs RDF ingestion, and if this is made available (can be done outside of Wikidata) then it would be easier to periodically sync datasets with Wikidata.

On that note however, you could export all Wikidata triples you need and just host that on your own SPARQL server (e.g. Jena) or use it with RDF tools like rdflib.

abraxaz··on Mindat.org, the largest open database of minerals, rocks, and meteorites
Can the database be exported somehow?
abraxaz··on Mindat.org, the largest open database of minerals, rocks, and meteorites
OWL for data model and RDF for data would work well for it, though I don't know if that is what they use.

Compare for example:

- Atelestite on mindat: https://www.mindat.org/min-407.html

- atelestite on WikiData (RDF based): https://www.wikidata.org/wiki/Q3627885

abraxaz··on Cue, an open-source data validation language
Thanks for the tip.
abraxaz··on Cue, an open-source data validation language
Can't see how, and this seems to suggest this functionality is not implemented yet: https://github.com/cuelang/cue/discussions/663

Mind sharing a reference?

abraxaz··on Cue, an open-source data validation language
Is there any plans to support model generation in future as can be done with JSON schema through something like https://github.com/quicktype/quicktype ?
abraxaz··on The data model behind Notion's flexibility
You don't have to use neo4j or any graph database to use RDF. It is just your current model seems very graph based and actually not that difficult to map to RDF, it would likely be possible to do with a jsonld context, and if you provided such a context then it would make your data a lot easier for others to consume.
abraxaz··on The data model behind Notion's flexibility
Is there a reason why you did not use rdf for representation and some rdf aware encoding like jsonld for serialization?

Would be significantly easier for others to work with, could easily query it with SPARQL.

abraxaz··on Legalese – Computational Law
You should submit this as a new link to HN, if you won't I will :)

Need more semantic tech related stuff on here IMO, most people working in tech don't even know it exists AFAICT.

abraxaz··on Legalese – Computational Law
Given that OWL and RDF is meant for both open world data and data models, and is amenable to partial specification, it really solves the problem of relating legal rules to the rest of the world quite well. In your legal ontology you just specify evidence, and admissible evidence, and then another ontology can pick up from there and model it in a specific domain.

And OWL has formal and provable entailment rules. So everyone can agree, given samme ontologies what the implications are. There are unsolved problems there but most of the things you list are already thought through.

abraxaz··on Semantic Finlex – Finnish Law and Justice as Linked Open Data
Some info about their SPARQL endpoints and example of queries: https://data.finlex.fi/en/sparql

This is still rudimentary I think, but clearly a step in the direction the world should be moving in.

abraxaz··on Legalese – Computational Law
Check this out also:

- https://data.finlex.fi/en/main - Semantic Finlex – Finnish Law and Justice as Linked Open Data

- https://lynx-project.eu/project - Legal Knowledge Graph for Multilingual Compliance Services

- https://www.mirelproject.eu/ - MINING AND REASONING WITH LEGAL TEXTS.

abraxaz··on Legalese – Computational Law
There has been a lot of work using OWL and RDF to represent laws and contracts and to use it for compliance checking.

For more information see:

- https://finregont.com/ : Semantic compliance in finance

- https://bankontology.com/ : Semantic bank compliance

- https://www.smartlogic.com/home : I think this company uses semantic technologies and AI to do legal compliance checking

- https://tel.archives-ouvertes.fr/tel-02062174/document : Automation of legal reasoning and decision based on ontologies (PhD thesis on the topic)

- https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=sema... : search google scholar for semantic legal compliance

I know there are more companies doing it, just don't know their names right now.

The benefits of this is:

- Don't need yet another DSL, RDF works fine, and it's already being widely used.

- A lot of the world is already modelled in OWL (e.g. Organizational Ontology, Provenance Ontology, literally too many to list)

- Already have a good query language with inference support (SPARQL)

abraxaz··on Google Dataset Search
Information on how to annotate datasets: https://developers.google.com/search/docs/data-types/dataset

> We can understand structured data in Web pages about datasets, using either schema.org Dataset markup, or equivalent structures represented in W3C's Data Catalog Vocabulary (DCAT) format. We also are exploring experimental support for structured data based on W3C CSVW, and expect to evolve and adapt our approach as best practices for dataset description emerge. For more information about our approach to dataset discovery, see Making it easier to discover datasets.

For more info on those:

- W3C's Data Catalog Vocabulary: https://www.w3.org/TR/vocab-dcat-3/

- Schema.org dataset: https://schema.org/Dataset

- CSVW Namespace Vocabulary Terms: https://www.w3.org/ns/csvw

- Generating RDF from Tabular Data on the Web (examples on how to use CSVW): https://www.w3.org/TR/csv2rdf/

abraxaz··on Pattern matching accepted for Python
> This while we have elephants in the room such as packaging. Researching best practices to move away from setup.py right now takes you down a rabbit hole of (excellent) blog posts, and yet you still need a setup.py shim to use editable installs, because the new model simply doesn't yet support this fundamental feature.

You can do editable installs with poetry, I do it every day.

Just run this: \rm -rv dist/; poetry build --format sdist && tar --wildcards -xvf dist/.tar.gz -O '/setup.py' > setup.py && pip3 install --prefix="${HOME}/.local/" --editable .

More details here: https://github.com/python-poetry/poetry/issues/34#issuecomme...

abraxaz··on Firefox 83 introduces HTTPS-Only Mode
Maybe try https://www.eff.org/https-everywhere then.
abraxaz··on Firefox 83 introduces HTTPS-Only Mode
Great to see this built into firefox, I have been using HTTPS Everywhere https://www.eff.org/https-everywhere to achieve similar results, it won't warn you if it is not https (i think) but it will try and upgrade to https if it can. It is available for chrome and firefox.

What particularly annoyed me was using http to sites which supported https.

abraxaz··on Mozilla Reaction to U.S. vs. Google
Yeah, AFAIK the donations all go to other causes. I could not find anything in writing regarding this on their site but this was explained to me on their IRC channel.
abraxaz··on Show HN: TidalWaves API – live, tokenized news metadata from around the world
Have you considered using JSON-LD/RDF for this? It would make it easier to integrate with WikiData and data from other sources.
abraxaz··on Catching use-after-move C++ bugs with Clang's consumed annotations
I think conan has real potential when it comes to package management

Some example of what can be done with it https://gitlab.com/xadix/xonan

It is kind of like gentoo portage or nix pkg and can be used to manage your tool chain also