GraphNews

6097 bookmarks
Custom sorting
helix-foundry: Your company's data, connected into one ontology, on your own computer. A local-first data workspace built on HelixDB and DuckDB.
helix-foundry: Your company's data, connected into one ontology, on your own computer. A local-first data workspace built on HelixDB and DuckDB.
·github.com·
helix-foundry: Your company's data, connected into one ontology, on your own computer. A local-first data workspace built on HelixDB and DuckDB.
Why AI Agents Forget: Graph Databases and the Agent Memory Problem
Why AI Agents Forget: Graph Databases and the Agent Memory Problem
Token Drop Episode 22: Arun Sharma of LadybugDB joins Sunil Baliga, Sajjad Khazipura, and Sam Pooni for a technical conversation on embedded columnar graph databases, MMAP, graph-vector hybrid search, temporal graphs, and the agent memory problem.
Why AI Agents Forget: Graph Databases and the Agent Memory Problem
·daax.ai·
Why AI Agents Forget: Graph Databases and the Agent Memory Problem
Enterprise semantic layers are largely middle ontologies
Enterprise semantic layers are largely middle ontologies
Enterprise semantic layers are largely middle ontologies. Yes, mapping to upper ontologies is important. But for the most part, semantic layer design deals in traditional questions of conceptual modeling, such as how to cope with legacy system schemas, how to represent business processes and business rules, how to integrate reference data, and how to create pragmatic, structured, and scalable systems for quotidian purposes like decision support. For a few years, I have worked through a variety of approaches to business process representation specifically. I have created process models in Prolog, in SQL, in Drools, in RDF, in Neo4j, in procedural and functional languages that support pattern matching, and in XML languages like XQuery. My first experiments used graph representations of state machines. But as I continued focusing on the problem, I realized the limits of the state machine approach: 1. We frequently need to remember variables with data types and even complex record types (customer records, order records, log files, telemetry data, etc.). State machines with their simplistic event labels don't excel at remembering variables. You can introduce stack machines, but you are still limited in expressivity and in typing discipline. 2. State machines suffer from state explosion. They can't model concurrency. They aren't flexible enough to use with modern ontology pipelines that leverage LLMs or NLP techniques to auto-construct knowledge graphs from both structured and unstructured data. 3. Validation and verification approaches for state machines leave much to be desired, specifically in the realm of measuring quality of business process models (more work has been done on state machine modeling for distributed software systems, but even there the inability of state machines to elegantly handle concurrency is a major limitation). For these reasons, I've recently moved my work on auto-generating knowledge graphs at scale onto colored Petri nets. Colored Petri nets are actually Turing complete (if you allow an infinite number of token colors). They are far more expressive than state machines, and easy to implement even in Neo4j. You can find established, algorithmic approaches to validation and verification that have been used across thousands of industrial organizations for decades. Data and record typing is consistent across transition guards, and there's a sensible way to understand how transitions consume typed objects and output (in some cases different) typed objects. Colored Petri nets have already been mapped to popular process modeling frameworks like SAP ARIS and OMG BPMN in order to provide formal semantics for these frameworks. So if your organization isn't already mandating an enterprise-wide framework for process modeling, why not simply implement the graph model that gives formal semantics to ALL the major process modeling frameworks directly in your database schemas? https://lnkd.in/gFtXqz46
Enterprise semantic layers are largely middle ontologies
·linkedin.com·
Enterprise semantic layers are largely middle ontologies
Decoupling Error Attribution in Cloud-Native Graph-RAG: A Data Integrity Diagnostic Framework
Decoupling Error Attribution in Cloud-Native Graph-RAG: A Data Integrity Diagnostic Framework
Graph-RAG systems often assume pristine data quality, overlooking the severe impact of perturbations in cloud-native databases. This paper proposes a three-layer decoupled diagnostic framework to...
Decoupling Error Attribution in Cloud-Native Graph-RAG: A Data Integrity Diagnostic Framework
·arxiv.org·
Decoupling Error Attribution in Cloud-Native Graph-RAG: A Data Integrity Diagnostic Framework
Fast thinking for knowledge graphs with Jev
Fast thinking for knowledge graphs with Jev
I have spoken several times about the slow/fast thinking analogy, how it applies to knowledge engineering and agentic memory. Thanks to Jev now everyone is talking about System 1 (fast thinking) models and this is precisely the emphasis I conveyed for a long time. How does Jev (or any universal classifier for that matter) help with knowledge graphs? There is entity typing and node classification. Once you've got a candidate mention/detection, deciding whether "...this is this a Person/Organization/Product/Legal Entity..." is a closed-vocabulary classification, not a generation task. Running that through a full LLM (say, Opus) is overkill and slow at scale. A Jev-style call with a typed schema over your ontology's classes is a much better fit. Jev will evaluate a state and returns typed answers and probabilities, which maps directly onto "...score this mention against my ontology's node types". Lots faster. Micro-seconds. You have entity resolution (ER) and dedup. ER pipelines spend most of their compute on pairwise "same entity or not" decisions. That's a binary-probability classification over a candidate pair, which is ideal for a cheap classifier as a first-pass filter before anything expensive (blocking, embedding similarity) runs on the survivors. Given the reported up to 200x faster inference and 400x lower cost than comparable LLMs on classification tasks, that's a meaningful cost cut if you're resolving entities. Triple validation. After an LLM (or a rule-based extractor) proposes a candidate triple, "...does this relation type actually hold given the context" is again a typed classification with a confidence score. Good for a provenance/confidence field on the triple rather than trusting the extractor's own reported certainty. There's triage before generation. This is probably the most useful pattern for your knowledge augmented AI stack. Use Jev as the fast layer that decides whether a case needs escalation to a full LLM at all and reserve the generative model for the residual hard cases. This is explicitly the recommended pattern from the ecosystem forming around it: use fast classifiers for most tasks and escalate uncertain cases to a reasoning model. Very much like the human fast/slow delegation process. Finally, guardrails on agentic graph construction. If you have an agent doing schema mapping, SPARQL/Cypher generation or write-actions against the graph, Jev fits as a middleware gate. A fast "...is this action safe/expected given current state" check before an LLM-driven agent commits a write. There are already Github projects taking advantage of this, see for instance JevGraph. I suspect many more will appear. JevGraph: https://lnkd.in/edUuqhbU #KnowledgeGraphs #Jev #AgenticMemory
Francois Vanderseypen
·linkedin.com·
Fast thinking for knowledge graphs with Jev
stop treating context as a buzzword, and start treating it as an engineerable data entity.
stop treating context as a buzzword, and start treating it as an engineerable data entity.
I disagree with a big group of friends and peers regarding what this years Big Data LDN was all about. I have been scrolling my feed and seen many posts about the shallowness of all the context buzz this year, but for me, that's not an important takeaway. Here is what I learned, and what I think we should be paying more attention to: Data: For me, there were two remarkable changes to what was once "the modern data stack". 1. Data Observability. Remember back when every single feature was a product of its own, in the modern data stack? I think the most extreme example was data lineage as a standalone product. I spoke to every CEO of those companies back when I published The Enterprise Data Catalog, and they were brilliant technologists!. But data lineage products dissappeared, either getting bought or successfully pivoting into economically viable solutions for data engineering. This year, at Big Data LDN, data observability are now seen as fully part of data governance platforms, and it confirms my reading of the market: data observability is becoming a feature too, not a product. 2. Data Catalogs: I know many have called this technology category impossible, unnecessary, and a bad idea. I conclude that this years Big Data LDN confirmed that there is no longer no way around a data catalog. Data catalogs are becoming for data, what CMDBs are for applications. As I see it, what happened the last 20 years, despite all the good intentions and generous investments, was that data architectures in every single industrial company has spun out of control. And that leaves data catalog in a weird, but prosperous state. One could say that many of them are becoming the victim of there own failure, but the fact is, they live on, _also_ as part of bigger platforms, as the vast majority of data catalogs are now bought. Context: I think we should stop treating context as a buzzword, and start treating it as an engineerable data entity. Together with Gaurav Patole, we had a 200 people standing room presentation (we expected 20 people), about this topic, at Big Data LDN. Basically, let's stop talking about context as a happenstance buzzword derived from context window and model context protocol. The task of capturing and serving context for agentic architectures is one of the most interesting and concretely challenging tasks these years, and successful technologists are already building, not pointing fingers at a word. I follow many that I cant mention/reveal because of my job, but one of the good ones to watch out for is Kepler brought to us all from my friend Vinoo Ganesh. **DISCLAIMER** The photo is from Mike Fergusons opening keynote, but my views are of course mine alone, I actually don't know what Mike thinks of my rant here. But the photo nicely captures how data catalogs live on, in a new era and new architecture. **DISCLAIMER** #data #observability #context #ai | 39 comments on LinkedIn
stop treating context as a buzzword, and start treating it as an engineerable data entity.
·linkedin.com·
stop treating context as a buzzword, and start treating it as an engineerable data entity.
Databricks And Palantir Built Ontologies For Two Jobs
Databricks And Palantir Built Ontologies For Two Jobs
Databricks And Palantir Built Ontologies For Two Jobs ⭕️ Databricks and Palantir now both sell an ontology for AI agents, built for two different jobs. Databricks' Genie Ontology tells an agent which definitions to trust when it answers [1]. The Palantir Ontology defines which changes an agent may make when it acts [2]. One request: count last quarter's active users across three user tables with no counting rule. Then merge the duplicates. 🔎 Rank the definitions, then filter them Genie Ontology turns existing tables, queries and dashboards into snippets of definition, such as "an active user is a distinct user, deduplicated across all platforms". It ranks each snippet by authority: its source, author, usage and freshness. Owner-reviewed definitions outrank inferred ones. Permissions constrain it: only snippets the asker may see reach the model, so answers differ by user. It infers and ranks instead of modelling every term by hand. ✍️ Type the change, then stage it Palantir models the merge as an action type, with its edits and rules. Each action needs an authorization grant. Because a merge changes live records, the agent by default only stages it in a scenario, a sandboxed copy, to review before committing. My sketch of the logic: action mergeusers(usera, user_b) allowed if caller holds the grant stage in a scenario commit after human review ⚖️ Two ways to let an agent write At Databricks the write goes through a tool, such as an MCP connection. At Palantir the write is an object in the ontology. A tool is faster to add. A typed action can be checked, granted and reviewed like data. I read this as two layers every agent stack needs: one for what the agent believes, one for what it may change. 🔭 Related work Research backs Databricks serving definitions at answer time. EvoOntology gave agents one semantic layer in two ways [3]. Pasted into the prompt, it scored below no layer at all, 64.6 against 69.5 averaged over four models (my arithmetic). Served as tools, it scored highest. In Ustimov's ladder, governed metrics gave the biggest jump, while a free-text knowledge base lowered accuracy [4]. Palantir's logic, a typed check before anything changes, must also cover what the agent writes. SHACL validation checks only the shapes it declares, so an invented term passes. In one benchmark, all 300 graphs with a fabricated term passed [5]. A closed-world check, which rejects any term the ontology never declared, caught every one. Both layers assume someone wrote the structure, and models struggle there. In CQ4OE, nine models recovered the terms of an ontology far more often than its hierarchy and relations [6]. EvoOntology's self-editing improved the ontology only behind a validation gate, and ended below its starting point without one. So the complete model has four parts. People own the structure, and agents propose edits behind a gate. Definitions are served as tools, and every write is checked against the ontology. | 38 comments on LinkedIn
Databricks And Palantir Built Ontologies For Two Jobs
·linkedin.com·
Databricks And Palantir Built Ontologies For Two Jobs
The research paper "ArchiMate Enterprise Hypermodel and Ontologies" submitted to MOVE 2024 is about to be published and made available in the proceeding of the conference after two long years.
The research paper "ArchiMate Enterprise Hypermodel and Ontologies" submitted to MOVE 2024 is about to be published and made available in the proceeding of the conference after two long years.
Congratulations on the publication, Nicolas... two years is a long wait but the topic is worth it. Your paper bridges ArchiMate and ontologies for interoperability. I've been working the same problem from the implementation side: ArchiMate 3.2 formalized as a native OWL/RDF ontology, with SHACL validation, derivation rules, and RDF-Star metadata. No bridge needed. The ArchiMate model *is* the knowledge graph. The blog series covering the design decisions is at https://lnkd.in/e4NetHbG
The research paper "ArchiMate Enterprise Hypermodel and Ontologies" submitted to MOVE 2024 is about to be published and made available in the proceeding of the conference after two long years.
·linkedin.com·
The research paper "ArchiMate Enterprise Hypermodel and Ontologies" submitted to MOVE 2024 is about to be published and made available in the proceeding of the conference after two long years.
Introducing edgextract: knowledge graph extraction that runs entirely in your browser. No server needed.
Introducing edgextract: knowledge graph extraction that runs entirely in your browser. No server needed.
Introducing edgextract: knowledge graph extraction that runs entirely in your browser. No server needed. Pulling knowledge graphs out of raw text has a dirty secret. Most pipelines ask a generative LLM to "extract all entities and relationships as JSON." A JSON schema makes sure the syntax is valid. It doesn't make sure the facts are true. Generative models invent relationships, flip link directions and give you no confidence score. On the CoNLL04 benchmark, a chat LLM produced 504 wrong relationships. So I tried a different approach. edgextract is an open-source Rust + WebAssembly engine. It builds knowledge graphs inside your browser tab using Tev1 0.8B, the small decision model from Together AI. 🧠 Stop generating. Start judging. edgextract doesn't ask the model to write hundreds of tokens of freeform JSON. It asks closed yes/no questions: → "Does Acme Inc use EdgeQuake?" 0.99 YES → "Did Acme Inc found PostgreSQL?" 0.01 NO The model can't make up new link types. Your ontology works as a firewall, so impossible pairs never get asked. Every relationship comes back with a probability score from 0 to 1. ⚡ Up to 18x faster → About 3 tokens per decision instead of hundreds of tokens of JSON → 77.8s with a local chat model writing JSON, 4.3s warm with Tev1 → Wrong links on CoNLL04 (zero-shot) cut from 504 to 218, with precision up 14 points 🔒 Private by design with WebGPU No API keys, no cloud and no token bill. Tev1 runs on your own GPU through ONNX Runtime Web. You can drop in confidential tech specs, M&A memos or clinical notes, and your data never leaves your machine. 🎚️ You set the cutoff Live sliders decide what gets into your graph: → 0.80 or higher: added automatically → 0.20 or lower: discarded → Anything in between: sent to a human review queue as plain-English sentences The result is an interactive D3 force graph where every link traces back to its source sentence and score. 📦 Available for Python and Rust, or as a web page with nothing to install. | 34 comments on LinkedIn
Introducing edgextract: knowledge graph extraction that runs entirely in your browser. No server needed.
·linkedin.com·
Introducing edgextract: knowledge graph extraction that runs entirely in your browser. No server needed.
7 types of agent memory
7 types of agent memory
7 types of agent memory (in 2 mins) 1. In-Context / Working Memory (Short-Term) Everything the model can currently see inside its context window: the system prompt, recent messages, tool outputs, and reasoning steps. 2. Semantic Memory (Long-Term) A persistent store of facts, preferences, and domain knowledge about a user or topic, decoupled from when it was learned. 3. Episodic Memory (Long-Term) A log of specific past events, full conversations, task runs, and what worked or failed, so the agent can learn from experience. 4. Procedural Memory (Long-Term) The agent's knowledge of how to do things: its skills, tool usage patterns, workflows, and behavioral rules. 5. External / Retrieval Memory (Short-Term + Long-Term) Knowledge stored outside the model in a vector database and pulled into context at inference time based on similarity search. 6. Parametric Memory (Long-Term) Knowledge baked directly into the model's weights during training: language, reasoning patterns, and general world knowledge. 7. Prospective Memory (Short-Term + Long-Term) The agent's ability to remember future intentions and scheduled goals, things it planned to do but has not yet executed, critical for long-horizon and multi-step planning agents. The practical takeaway: start with working memory. Add semantic memory when users expect the agent to remember them across sessions. Layer in episodic, procedural, and prospective memory only when your agent needs to plan ahead, learn from failure, and adapt over time. Knowing what each one does will help you build the right system for the right problem. System designing is very important when building memory! -- ♻️ Repost if you found it helpful! ➕ We often this discuss real-world AI/ML here 👇 ➕ Join 48.000+ AI/ML builders here: https://lnkd.in/ds_SzEUH | 62 comments on LinkedIn
7 types of agent memory
·linkedin.com·
7 types of agent memory