GraphNews

6069 bookmarks
Custom sorting
The Semantic Units Framework, a technology-agnostic representational approach to FAIR and CLEAR knowledge infrastructures - Scientific Data
The Semantic Units Framework, a technology-agnostic representational approach to FAIR and CLEAR knowledge infrastructures - Scientific Data

The Semantic Units Framework by Lars Vogt is a spiritual follow-up to Rosetta Statements, and a far more detailed exploration of the wider idea. To recap, Rosetta Statements were more of a tool for converting simple English language statements (“n-ary” statements, you might recall) into RDF triples. The idea was to encode some basic rules of English syntax into RDF to make knowledge graph construction easier.

The broader Semantic Units Framework, however, is broader in scope, and more abstract. Vogt refers to it as ‘conceptual in nature’ and more of a philosophical approach to knowledge graph creation, rather than a proposal for any specific technology or standard. It can be applied to RDF (and indeed they use an RDF framing within the paper) but it is not specific to RDF. In fact, a lack of universal meaning to the term ‘graph technology’ is cited as one of the many motivations for the paper.

Rather, the paper talks more generally about how natural language can be used to construct FAIR knowledge graphs, and various challenges that may be encountered. There is a particular emphasis on how to expand the notion of knowledge and information beyond that of ‘individual facts’, which is how a single graph entity (or RDF triple) is often constructed. How do we handle statements that may be only mostly or partially true? How do we handle abstract, existential claims, or context?

Vogt proposes a two fold approach: (1) the quantification of ‘semantic units’ as first-class objects, and (2) the introduction of a set of logical resource categories to model difference types of claim. The discussion is lengthy, with an extensive list of examples. It will be interesting to if his suggestions are widely adopted as an increasing number of systems struggle to keep up with the volume of natural language data now available.

h/t G.V()

·nature.com·
The Semantic Units Framework, a technology-agnostic representational approach to FAIR and CLEAR knowledge infrastructures - Scientific Data
Knowledge graphs in drug discovery: How network biology is finding new targets | Drug Discovery News
Knowledge graphs in drug discovery: How network biology is finding new targets | Drug Discovery News

Ever wondered how the biomedical industry uses knowledge graphs? Here’s a general introductory guide from Drug Discovery News.

Did you know that the vast majority of modern pharmaceuticals conceptually do the exact same thing? Nearly all prescribed drugs are just chemicals designed to modulate the behavior of proteins. In many ways, “thing that modulates a protein” is a good approximate definition of the word “drug”.

If that makes biomedical research sound simple however, you’re out of luck. The problem is the sheer size of the parameter space. There are estimated ~20,000 protein-coding genes in the human body, and an even higher number of possible diseases, drugs, phenotypes and pathways. Keeping on top of all the different possible combinations and relationships between them by hand would be impossible, but is fortunately the kind of thing that a graph database can do very well.

As this introduction mentions, though, reality is even more complicated than that. Even “the graph is a simplification of biology.” Just because we give an entry and a name (perhaps even a chemical formula) to a protein, drug or disease, that doesn’t mean that its properties and behaviours can be fully predicted from base principles or existing data. Biomedical research has many strongly supported relationships, but also relationships that are merely suggested, inferred, or perhaps even pure hypothesis. This incompleteness or uncertainty in knowledge also has to be captured.

This article refers to major biomedical knowledge graphs: PrimeKG, Open Targets, Hetionet, STRING, and mentions how neural networks can be used look for patterns in the complicated space of biomedical processes.

h/t weekly edge G.V()

·drugdiscoverynews.com·
Knowledge graphs in drug discovery: How network biology is finding new targets | Drug Discovery News
What Palantir Got Right With Ontologies - Scrydon
What Palantir Got Right With Ontologies - Scrydon
Palantir put the ontology — not the LLM — at the centre of enterprise AI, and that architectural bet is being proven right. But AI itself has now collapsed the cost of building an ontology. You no longer need an army of forward-deployed engineers, and you certainly don't need to hand your most sensitive data to a US vendor to get one.
·scrydon.com·
What Palantir Got Right With Ontologies - Scrydon
A few weeks ago everyone was talking about loops. Now it's graphs.
A few weeks ago everyone was talking about loops. Now it's graphs.
A few weeks ago everyone was talking about loops. Now it's graphs. Both live or die on one thing: your company brain. This week, Shann Holmberg (@shannholmberg on X) shared how he runs his marketing on graphs. Here's the difference (resolving a support ticket): 𝗟𝗼𝗼𝗽𝘀 → you set the frame, the agent owns the path. Hand it a goal and a bar. It drafts, checks, fixes, and loops until it clears. You don't pick the steps. The agent does. 𝗚𝗿𝗮𝗽𝗵𝘀 → you draw the path first, the agent fills each node. read → triage → draft → QA → send, with a checkpoint at each step. The agent solves each box, but the map is yours. Every one of those nodes pulls from the same place. Triage needs past tickets. Draft needs the policy and product docs. That shared context (your company brain) is what each box reaches into as the work moves through. The brain has that map, the graph is that process, inferred explicitly. 𝗦𝗼 𝘄𝗵𝗲𝗻 𝗱𝗼 𝘆𝗼𝘂 𝘂𝘀𝗲 𝘄𝗵𝗶𝗰𝗵? → Reach for a 𝗹𝗼𝗼𝗽 when it's one-off work, no clear path yet. Let the agent figure it out. → Reach for a 𝗴𝗿𝗮𝗽𝗵 when it's repeatable work you already know the steps for. Lock them in. Loops are for exploring. Graphs are for scaling. Which one are you running right now? 👀 | 26 comments on LinkedIn
A few weeks ago everyone was talking about loops. Now it's graphs.
·linkedin.com·
A few weeks ago everyone was talking about loops. Now it's graphs.
All Relations Lead to Rome: Automated Knowledge Graph Creation and Question Generation
All Relations Lead to Rome: Automated Knowledge Graph Creation and Question Generation

All Relations Lead to Rome (ARLtR), from Matthijs Jansen op de Haar (Twente), Tobias Stähle (ETH Zürich) and Lorenzo Gatti (Twente), is a new release aiming to serve as a new benchmark for information retrieval.

The researchers claim that existing datasets exist largely in two exclusive types: (1) vector-based retrieval over unstructured text or (2) reasoning over a knowledge graph. They argue that there are few existing benchmarks that combine both in one place.

Their work, ARLtR, aims to be just such a unified dataset. It offers everything in one: knowledge graph, embeddings, and question and answer pairs explicitly grounded in the entities, relations, and supporting text sourced from a central corpus of documents. The name isn’t just metaphorical, the benchmark (comprising 19,000 entities, 16,000 chunks, and 8,400 question/answer pairs) is literally a collection of data concerning the Roman Empire.

The idea is that coupling the symbolic graph with the dense vectors can give a single coherent resource for evaluating and developing hybrid retrieval systems and “semantic steering” approaches. This is a dataset that might be useful to those building GraphRAG-style systems, and it’s on Hugging Face if you want to check it out.

https://huggingface.co/datasets/FaynePro/all-relations-lead-to-rome

·arxiv.org·
All Relations Lead to Rome: Automated Knowledge Graph Creation and Question Generation
OKF v0.2 is here. A few weeks back we put out the Open Knowledge Format (OKF): a standard for getting the context agents need out of proprietary APIs and into plain markdown and YAML. What we didn't expect was how fast the conversation moved to what we're shipping (and had planned to ship) in v0.2: trust.
OKF v0.2 is here. A few weeks back we put out the Open Knowledge Format (OKF): a standard for getting the context agents need out of proprietary APIs and into plain markdown and YAML. What we didn't expect was how fast the conversation moved to what we're shipping (and had planned to ship) in v0.2: trust.
OKF v0.2 is here. A few weeks back we put out the Open Knowledge Format (OKF): a standard for getting the context agents need out of proprietary APIs and into plain markdown and YAML. What we didn't expect was how fast the conversation moved to what we're shipping (and had planned to ship) in v0.2: trust.
·linkedin.com·
OKF v0.2 is here. A few weeks back we put out the Open Knowledge Format (OKF): a standard for getting the context agents need out of proprietary APIs and into plain markdown and YAML. What we didn't expect was how fast the conversation moved to what we're shipping (and had planned to ship) in v0.2: trust.
Live now: the full AI x Graphs Track from AI Engineer World's Fair, in partnership with Neo4j.
Live now: the full AI x Graphs Track from AI Engineer World's Fair, in partnership with Neo4j.
Live now: the full AI x Graphs Track from AI Engineer World's Fair, in partnership with Neo4j. Watch it here: https://lnkd.in/eNGV6kdQ The hard problem in agent engineering is not only making models smarter. It's giving them durable, inspectable context: what happened, what is true, where a fact came from, and how the pieces connect. Emil Eifrem, CEO of Neo4j, makes the case for thinner agents on top of a shared semantic substrate: an ontology of the business, the systems behind it, and the execution traces that let every agent learn from the last one. Yohei Nakajima, creator of BabyAGI, flips the usual architecture around. In ActiveGraph, the event log is the agent: state becomes a graph, and replays, rollbacks, forks, and controlled self-improvement follow from that design. Daniel Chalef of Zep AI tackles the provenance problem. When an LLM synthesizes facts from multiple sources, a source ID is not enough. In Graphiti, provenance is itself a graph—so agents can trace claims, apply trust policies, and delete data without losing the audit trail. Also on the track: - Zach Blumenfeld, Neo4j - James Le, TwelveLabs - frank coyle, UC Berkeley - Mike Phipps, Gates Foundation - Ritvik Pandya, JPMorgan Chase - Omri Bruchim & Tomer Ast, monday[.]com - Stephen Chin, Neo4j - Shafik Q. & Joanne Song, The New York Times - Subbiah Sethuraman & Abhilash Asokan, ZS Associates From graph memory and video context to agentic constraints, data models, and knowledge graphs as a control plane: this is a deep look at what it takes to give agents context they can actually reason over. Full track: https://lnkd.in/eNGV6kdQ
Live now: the full AI x Graphs Track from AI Engineer World's Fair, in partnership with Neo4j.
·linkedin.com·
Live now: the full AI x Graphs Track from AI Engineer World's Fair, in partnership with Neo4j.
Existing Postgres extension goes through DuckDB.
Existing Postgres extension goes through DuckDB.
Existing Postgres extension goes through DuckDB. A new extension tentatively named pg_client will use postgres client libraries. A postgres extension (pg_ladybug - good guess!) will alllow you to use ladybug as the cypher engine instead of the graph query capabilities in Apache AGE or Postgres 19.
Existing Postgres extension goes through DuckDB.
·linkedin.com·
Existing Postgres extension goes through DuckDB.
When I describe what I do as a taxonomist to designers, I have heard the response, “that sounds like a design system”
When I describe what I do as a taxonomist to designers, I have heard the response, “that sounds like a design system”
When I describe what I do as a taxonomist to designers, I have heard the response, “that sounds like a design system” a couple of times now. I had a superficial understanding of design systems—I knew they were essentially controlled vocabularies for UI components that evolved somewhat convergently with semantic modeling. However, I couldn’t speak on the subject beyond that. Curious, I decided to dig into design systems from a taxonomist lens to see how similar they really are. My partner is a designer, so I started by asking him about the reality of working with design systems, and realized governance is a pain point we have in common. Without dedicated resourcing, mounting technical debt is inevitable. For design systems, it manifests as ad-hoc components and hardcoded values; for semantic taxonomies, it’s semantic drift and tag bloating. After sharing his governance struggles, he pointed me to open-source systems as examples of what managed design looks like at scale. I saw how uniquely varied these open source systems can be, especially in their naming conventions. In semantic modeling, there are fairly established best practices for naming concepts. But in design systems, I noticed distinct philosophies. The Atlassian Design System uses granular and descriptive token names—you know how the token fits in the schema by reading it. Adobe Spectrum, on the other hand, uses leaner token names that rely heavily on rich external documentation. As a taxonomist, the former seems very familiar, but I don't think there's an analog in taxonomy for the latter. Naming conventions are important because in many design systems, the entire schema and hierarchy are encoded directly inside the token name string itself. Because relationships are bound to a delimited string, most token architectures are forced into strict taxonomic hierarchies. However, real UI logic—handling dark mode, multi-brand themes, accessibility overrides, and state variations—is inherently multidimensional. In semantic modeling, this is a classic use case for a graph. Instead of modeling just strict broader and narrower tree structures, we could capture these multidimensional rules through custom, contextual relationships in a web or graph. I can’t help but envision a graph-based design system that scales as a product ecosystem grows. I’d love to hear from designers and information architects alike – what structural challenges are you facing that a graph might solve? #DesignSystems #Taxonomy #SemanticModeling #InformationArchitecture #UXDesign Photo by Mike Baumeister on Unsplash
When I describe what I do as a taxonomist to designers, I have heard the response, “that sounds like a design system”
·linkedin.com·
When I describe what I do as a taxonomist to designers, I have heard the response, “that sounds like a design system”
a growing awareness that we are seeing a profound shift in which graphs become the substrate of computation for AI, rather than loops (which are imperative structures).
a growing awareness that we are seeing a profound shift in which graphs become the substrate of computation for AI, rather than loops (which are imperative structures).
I spent the day at Graphcon 27 in Seattle. While I will cover the convention in a subsequent post, there was a growing awareness that we are seeing a profound shift in which graphs become the substrate of computation for AI, rather than loops (which are imperative structures). This changes things dramatically, because it forces a declarative mode of thinking - more iterators, more recursion, with the role of a control plane for the most part being used to set the cursor - where you are in which graph - as a focusing agent. This necessitates that you think about problems differently. A signal is added as a small graph within a larger one, the signal in turn establishing a focus that can then be tested within a given context. If the signal satisfies that context, a transformation generates a new subgraph within the broader graph, embedding this subgraph as new, perhaps partial information. You never replace content, you only create new records that can be identified as the head of a chain of previous records. This creates memories. Queries can either retrieve linear patterns (tables of vectors) that map cleanly to the more relational paradigm, or can create new subgraphs that can be placed in different named graphs, not always within the same graph store. Periodically, as with any data store, you have to persist the state of the graph in a more durable media (documents, for instance), then compact and distill the relevant information to create new "ground truths". This was what I came away with today - many different researchers, from very different backgrounds, working on different forms of graphs, each coming to the same pattern. The control plane establishes focus - which node in which graph are you working on now, which can then be handed off to the AI agent to establish context and update their state predictably, persistently, and consistently. Far from making graphs irrelevant, graphs are taming the exuberant flights of fantasy of the LLM in order to harness that animus for constructive work, but its going to change the nature of software engineering profoundly as a consequence. https://lnkd.in/gJUku5ez
a growing awareness that we are seeing a profound shift in which graphs become the substrate of computation for AI, rather than loops (which are imperative structures).
·linkedin.com·
a growing awareness that we are seeing a profound shift in which graphs become the substrate of computation for AI, rather than loops (which are imperative structures).
Choosing the right ontology reasoner depends on your ontology's complexity and your use case, not every project needs the same one.
Choosing the right ontology reasoner depends on your ontology's complexity and your use case, not every project needs the same one.
New to Ontology? Ever wondered what all these reasoners do and why we have so many? When you open the Reasoner menu in Protege, you'll see several options. Here's a quick overview: ELK – Super fast reasoner for the OWL 2 EL profile. Great for very large biomedical ontologies like SNOMED CT, but it doesn't support the full expressiveness of OWL. HermiT – A complete OWL 2 reasoner. Excellent for checking ontology consistency, class hierarchies, and complex logical inferences. Slower than ELK but much more expressive. Ontop – Not a traditional ontology reasoner. It enables Ontology-Based Data Access (OBDA) by mapping relational databases to an ontology, allowing you to query databases using SPARQL. Pellet – A full-featured OWL reasoner supporting consistency checking, classification, realization, and datatype reasoning. Also supports SWRL rules. Pellet (Incremental) – Optimized for ontologies that change frequently. Instead of reasoning over the entire ontology after every edit, it updates only the affected parts, making repeated reasoning faster. Which one should you use? Speed on large EL ontologies - ELK Full OWL reasoning - HermiT Query relational databases through an ontology - Ontop Need SWRL support or Pellet-specific features - Pellet Frequent ontology edits - Pellet (Incremental) Choosing the right reasoner depends on your ontology's complexity and your use case, not every project needs the same one. #Ontology #SemanticWeb #KnowledgeGraphs #OWL #Protégé #Reasoning #RDF #LinkedData
Choosing the right reasoner depends on your ontology's complexity and your use case, not every project needs the same one.
·linkedin.com·
Choosing the right ontology reasoner depends on your ontology's complexity and your use case, not every project needs the same one.
From Loops to Graphs | LinkedIn
From Loops to Graphs | LinkedIn
Three editions ago the rung was the loop; last week the discourse climbed again and named the graph. This week’s focus introduces graph engineering, what the weeks-old term actually means, and reads the current state of the software factory debate through it, from OpenAI’s million-line no-human-code
·linkedin.com·
From Loops to Graphs | LinkedIn
Knowledge augmented AI part 1. You can store facts, and you can query facts, but somewhere you also need to derive facts and that's a different kind of computation entirely. It happens in the materialization (aka projection). This is where Datalog quietly does more work than it gets credit for.
Knowledge augmented AI part 1. You can store facts, and you can query facts, but somewhere you also need to derive facts and that's a different kind of computation entirely. It happens in the materialization (aka projection). This is where Datalog quietly does more work than it gets credit for.
Knowledge augmented AI (KAAI) is the new all-embracing knowledge graph architecture combining a RDF semantic layer and an LPG operational layer. It does not fit in a single LinkedIn post, so I will sketch it one concept at a time. You can store facts, and you can query facts, but somewhere you also need to derive facts and that's a different kind of computation entirely. It happens in the materialization (aka projection). This is where Datalog quietly does more work than it gets credit for. You give it rules, it applies them, checks if anything new emerged, and repeats until the derivation stabilizes and stops producing anything different. No procedural loop, no explicit recursion depth, just: iterate the rule set over the fact set until it converges. That's the entire semantics. The three-layer pattern most of us are converging on (ingestion, semantic, operational) has a gap in the middle. The semantic layer defines what's true in principle (a supplier trust chain, a consent scope, a hierarchy of access). The operational layer needs answers right now, as graph queries an application can call. Datalog is the bridge: it takes principled rules and materializes or streams them into facts the operational layer can actually traverse. Transitive consent, degrading trust across a supply chain, bitemporal validity windows and so on. These aren't queries, they're derivations, and derivations want fixpoint semantics, not another JOIN. The platform landscape here is more alive than people assume. RDFox does OWL 2 RL reasoning as compiled Datalog and is genuinely fast at it. Vadalog (Oxford) pushes into existential rules for exactly this kind of enterprise ontological reasoning. Soufflé compiles Datalog to native code for static-analysis-scale workloads. CozoDB (not Kuzu but Cozo) embeds Datalog directly as a query language over a graph/relational store, which is a nice preview of where operational layers might head. Ontotext's GraphDB and Stardog both fold rule materialization into their reasoning layers rather than treating it as a separate step. The pattern worth noticing: none of the popular LPG engines speak Datalog natively. That absence is precisely why the reasoning layer keeps reappearing as its own architectural component rather than disappearing into "just another index." And once those derived facts exist, they still need a surface a human can actually look at and trust which is its own layer of work, and a good argument for keeping visualization (yFiles, Ogma) as a first-class citizen rather than an afterthought bolted onto the query result. Datalog isn't the newest idea in this stack but the one holding the middle together. ▶ RDFox: https://lnkd.in/e4hz-gJ8 ▶ Vadalog: https://lnkd.in/ekWtwx3e ▶ GraphDB: https://lnkd.in/eXF_p_XE ▶ CozoDB: https://www.cozodb.org/ ▶ Soufflé: https://lnkd.in/eGYwZFa2 ▶ Stardog: https://www.stardog.com/ #KnowledgeGraphs #RDF #Datalog #GraphDatabases | 14 comments on LinkedIn
Knowledge augmented AI (KAAI) is the new all-embracing knowledge graph architecture combining a RDF semantic layer and an LPG operational layer. It does not fit in a single LinkedIn post, so I will sketch it one concept at a time.You can store facts, and you can query facts, but somewhere you also need to derive facts and that's a different kind of computation entirely. It happens in the materialization (aka projection). This is where Datalog quietly does more work than it gets credit for.
·linkedin.com·
Knowledge augmented AI part 1. You can store facts, and you can query facts, but somewhere you also need to derive facts and that's a different kind of computation entirely. It happens in the materialization (aka projection). This is where Datalog quietly does more work than it gets credit for.
What We Learned Building Enterprise MCP at Bloomberg | Without rich semantic metadata, AI systems can't reason reliably over financial data.
What We Learned Building Enterprise MCP at Bloomberg | Without rich semantic metadata, AI systems can't reason reliably over financial data.
Model Context Protocol (MCP) is a genuinely important standard, a connection layer that gives AI systems structured access to external tools and data sources at scale. But connection and readiness aren’t the same thing, and financial institutions building AI strategies on top of early-stage MCP offe
seman
·linkedin.com·
What We Learned Building Enterprise MCP at Bloomberg | Without rich semantic metadata, AI systems can't reason reliably over financial data.
Knowledge Graph vs. Vector Database: Choosing your enterprise AI foundation
Knowledge Graph vs. Vector Database: Choosing your enterprise AI foundation
🟥Knowledge Graph vs. Vector Database: Choosing your enterprise AI foundation.   Too many enterprise AI initiatives stall because teams treat their underlying data layer as a single, uniform commodity.   If you are building intelligent corporate systems, you aren't just storing text strings, you are mapping organizational memory. And the way you store that memory dictates how confidently your models can reason.   Under the architectural microscope, the two primary data pillars serve entirely different cognitive functions:       🟥 Vector Databases (Dense Semantic Retrieval)   • Best for: Rapid semantic similarity, fuzzy matching, and searching across massive, unstructured text blocks.   • How it works: It converts data into mathematical coordinates. It excels at finding things that "sound similar" contextually.   • The Limit: It is structurally blind. A vector store knows two concepts are closely related in a document, but it cannot tell you the explicit, deterministic logic of how they connect.       🟥 Knowledge Graphs (Explicit Relational Structure)   • Best for: Strict entity mapping, deterministic business rules, and multi-hop data traversal.   •How it works: It links data points as explicit Nodes (entities) and Edges (relationships).   •The Limit: High upfront layout cost. It relies heavily on accurate initial entity extraction and defined schemas to function effectively at scale.   🟥 The Hybrid Reality:   The most resilient enterprise intelligence engines don't pick a side. They run a dual-engine architecture.   By running a hybrid pipeline, vectors handle the initial semantic capture across your documentation, while a knowledge graph acts as the strict logical anchor. The graph ensures the LLM respects the hierarchy of the data, virtually eliminating the structural hallucinations that plague standard RAG setups.   Are you building a pure vector play to get to market quickly, or are you already introducing graph topology to govern your LLM context windows?   #dataarchitecture #aiengineering #knowledgegraphs #vectordb #rag
Knowledge Graph vs. Vector Database: Choosing your enterprise AI foundation
·linkedin.com·
Knowledge Graph vs. Vector Database: Choosing your enterprise AI foundation