The Semantic Units Framework by Lars Vogt is a spiritual follow-up to Rosetta Statements, and a far more detailed exploration of the wider idea. To recap, Rosetta Statements were more of a tool for converting simple English language statements (“n-ary” statements, you might recall) into RDF triples. The idea was to encode some basic rules of English syntax into RDF to make knowledge graph construction easier.
The broader Semantic Units Framework, however, is broader in scope, and more abstract. Vogt refers to it as ‘conceptual in nature’ and more of a philosophical approach to knowledge graph creation, rather than a proposal for any specific technology or standard. It can be applied to RDF (and indeed they use an RDF framing within the paper) but it is not specific to RDF. In fact, a lack of universal meaning to the term ‘graph technology’ is cited as one of the many motivations for the paper.
Rather, the paper talks more generally about how natural language can be used to construct FAIR knowledge graphs, and various challenges that may be encountered. There is a particular emphasis on how to expand the notion of knowledge and information beyond that of ‘individual facts’, which is how a single graph entity (or RDF triple) is often constructed. How do we handle statements that may be only mostly or partially true? How do we handle abstract, existential claims, or context?
Vogt proposes a two fold approach: (1) the quantification of ‘semantic units’ as first-class objects, and (2) the introduction of a set of logical resource categories to model difference types of claim. The discussion is lengthy, with an extensive list of examples. It will be interesting to if his suggestions are widely adopted as an increasing number of systems struggle to keep up with the volume of natural language data now available.
h/t G.V()
Ever wondered how the biomedical industry uses knowledge graphs? Here’s a general introductory guide from Drug Discovery News.
Did you know that the vast majority of modern pharmaceuticals conceptually do the exact same thing? Nearly all prescribed drugs are just chemicals designed to modulate the behavior of proteins. In many ways, “thing that modulates a protein” is a good approximate definition of the word “drug”.
If that makes biomedical research sound simple however, you’re out of luck. The problem is the sheer size of the parameter space. There are estimated ~20,000 protein-coding genes in the human body, and an even higher number of possible diseases, drugs, phenotypes and pathways. Keeping on top of all the different possible combinations and relationships between them by hand would be impossible, but is fortunately the kind of thing that a graph database can do very well.
As this introduction mentions, though, reality is even more complicated than that. Even “the graph is a simplification of biology.” Just because we give an entry and a name (perhaps even a chemical formula) to a protein, drug or disease, that doesn’t mean that its properties and behaviours can be fully predicted from base principles or existing data. Biomedical research has many strongly supported relationships, but also relationships that are merely suggested, inferred, or perhaps even pure hypothesis. This incompleteness or uncertainty in knowledge also has to be captured.
This article refers to major biomedical knowledge graphs: PrimeKG, Open Targets, Hetionet, STRING, and mentions how neural networks can be used look for patterns in the complicated space of biomedical processes.
h/t weekly edge G.V()
All Relations Lead to Rome (ARLtR), from Matthijs Jansen op de Haar (Twente), Tobias Stähle (ETH Zürich) and Lorenzo Gatti (Twente), is a new release aiming to serve as a new benchmark for information retrieval.
The researchers claim that existing datasets exist largely in two exclusive types: (1) vector-based retrieval over unstructured text or (2) reasoning over a knowledge graph. They argue that there are few existing benchmarks that combine both in one place.
Their work, ARLtR, aims to be just such a unified dataset. It offers everything in one: knowledge graph, embeddings, and question and answer pairs explicitly grounded in the entities, relations, and supporting text sourced from a central corpus of documents. The name isn’t just metaphorical, the benchmark (comprising 19,000 entities, 16,000 chunks, and 8,400 question/answer pairs) is literally a collection of data concerning the Roman Empire.
The idea is that coupling the symbolic graph with the dense vectors can give a single coherent resource for evaluating and developing hybrid retrieval systems and “semantic steering” approaches. This is a dataset that might be useful to those building GraphRAG-style systems, and it’s on Hugging Face if you want to check it out.
https://huggingface.co/datasets/FaynePro/all-relations-lead-to-rome