2008-09-26
Open GUID : anchoring hubjects
Jason Borro has announced yesterday his Open GUID initiative on the Linking Open Data forum. After a first day of open discussion, it appears that he has came with the right implementation of hubjects, and moreover with a great metaphor. Hubjects must be anchored in signs, bot human-readable and computer-readable. Here is what I come with this morning.
2008-07-15
Everything is a Sign
That is, every thing is a sign. The first and main function of any language is to allow division of the world into "this" and "not this", based on some interpretation of data received from the world. Such an interpretation of data as signs is the basic form of semiosis, a process performed for quite a while by humans, and for many more ages by animals before them. It can now be performed by machines or information systems (roughly, computers connected to data acquisition devices). The aspects of this process can be defined as following.
- SALIENCE : Capacity to separate as meaningful (significant or salient) a certain data set from the continuous data flow we get from the world through our perceptive experience, be it direct through our biological senses, or indirect through one or more several levels of mediation : reading data gathered by instrumental devices, compilation of such data over time, texts interpreting those data.
- SIGNIFICATION : Capacity to consider the salient data set as a signifier conveying a particular meaning (signified), based on some characteristics such as spatial connectivity, permanence in time, regularity of patterns, similarity with other data sets previously interpreted and stored as signs, or anything the interpreter sees fit by its own rules and general view of the world. The core and essential meaning assigned is generally permanence, existence of a "thing" underlying the "sign". The thing is the signified associated to the signifier which is the data set.
- REPRESENTATION : Translate this sign/thing (both signifier and signified) into some proxy in a representation language allowing storage and retrieval for further use. Typical forms of representation include assignation of identifiers (symbols, icons, names, code numbers), description of the signified, and its connection to pre-existing ones through classification, typing, or any other kind of association or linking.
The above analysis can be set as the basis for a general semiosis framework applicable to natural languages (human or otherwise), formal languages used in our information systems, and scientific languages (theories in physics, biology). This framework, while keeping agnostic at the metaphysical level on the ontological status of things, will hopefully help to provide a solid theoretical foundation to the emerging semiosphere, the network of human knowledge and languages and information systems.
For use of this approch in the Semantic Web area, see a first cut ontology here.
[Note 2013-02-05] : This post has been for years and is still the top viewed in this blog, and I really don't understand why. Passer-by if you care to tell me how you came here, please comment below. Thanks!
2008-05-21
Own your identity
ownyouridentity.com is a blog about owning your online identity in a world with an increasing amount of software that wants to own it for you. It's written by the chi.mp team and friends.
2008-05-18
Managing Co-reference (Was: A Semantic Elephant?)
One more public thread on SW list on our favourite issue. But some interesting points to note in the current discussion, well summed up by Aldo Gangemi at mid-course.
- We definitely need some property or mechanism, weaker than owl:sameAs, to assert that two URIs have similar referents.
- The semiotic aspects of co-reference are more and more acknowledged, even by formal logic gurus.
2007-10-09
URIs for languages
I've eventually given up trying representing hubjects at all, at least for the moment. I had a serious try at it at lingvoj.org. But after discussions in the Linking Open Data forum, I eventually surrendered and published the languages description in a way conformant to W3C recommandations for Semantic Web architecture, with content negociation, 303 redirects and the like. I've even suppressed the previous post here saying otherwise, which would be now full of dead links and would bring about confusion.
So we'll see how this flies. Feed your favourite tool with the URI http://www.lingvoj.org/lang/zh, and figure by yourself if it provides a useful description of the Chinese language, both for humans and machines.
2007-07-24
Using owl:sameAs in Linked Data
It's been a very long and interesting thread on Linking Open Data forum and elsewhere, about the use and semantics of owl:sameAs. I just suggested the following best practices :
- Assertions such as "a:foo owl:sameAs b:bar" should be grounded on some form of agreement of the owners of a:foo and b:bar, on whichever basis they both decide to agree.
- For outsiders (owning neither a: or b: domains), such agreement could be shown by the presence of the assertion in symmetrical way in both domains, each domain using its own URI/resource on subject side, and the other's on object side, that is :
(a) asserts "a:foo owl:sameAs b:bar"
(b) asserts "b:bar owl:sameAs a:foo". - If one side (a) pushes the assertion first, the other side (b) should be at least made aware of it by (a), and is entitled to say she agrees or not : (a) says that "a:foo owl:sameAs b:bar", but as the owner of (b), I do not necessarily agree. Such lack of agreement could be implicitly entailed from the absence of the reciprocal assertion on (b) side.
2007-04-24
A journey to Data Mountains
Seems about time to revisit a famous Zen aphorism (more verbose translation here)
Morning mountains, only mountainsNoon mountains, more than mountainsEvening mountains, simply mountains
How does this apply to what we have been speaking about here? We've tried in the morning of our ignorance to pile up mountains of data, and had hard time making sense of them.
In the broad daylight of our powerful abstract thought, both intuition and logic, we found out that useful data were data about some thing(s). Making explicit the things data are about was really the way to follow to organize, understand, search, query, in a thousand ways make data more useful, more meaningful. So we went through those mountains with this "about-ness" in mind, and they looked indeed more than mountains, they looked like information about things. We called it classes, properties, relations, and started re-engineering data in all sorts of smart ways : metadata, RDF, topic maps, ontologies and the like. Happy to bring more and more meaning into data, we called it knowledge, and we thought we had found answers to some fundamental questions, discovered the information lost in data, and the knowledge lost in information. Captured the mountain's spirit.
Then came the time to harvest, to weave it all together. I had knowledge in my information system, and so did you. Or so we figured. But looking into your system, I found only data, and so did you when you looked into mine. Where are the things gone? Where is the knowledge hidden? Only data, which we had to figure again together how to weave.
So we'd not captured the mountain's spirit after all. But we did not travel in vain, because we've felt it blowing by. We know it's hidden somewhere beyond the data, beyond information, beyond what we have called knowledge, and that without this spirit our data would be completely meaningless and useless indeed. And what is more, we have at least found a piece of wisdom lost in knowledge : data are only data.
Subscribe to:
Posts (Atom)