2010-04-08

Coreference as a Service

Yahoo! releases Concordance as part of GeoPlanet API. The aim of this service is to provide equivalence between identifiers for geo entities defined in different namespaces. Quoting Gary Gale on Yahoo! Geo Technologies Blog.

We’ve collected these identifiers and namespaces as a single object, a concordance, which empowers a user to reference each source. You can think of it as a mapping of an identifier in a namespace to its equivalent in another namespace. But it’s not a joining of information; we’re only enumerating the identifiers, not the back-end data or attributes that they describe.
The last sentence is important. The service is agnostic on the data model or ontologies used by the various identifiers publishers. Ontological emptiness makes the service useful.

Another striking example is provided by Ellerdale, reconciliating Wikipedia or Freebase topics with Twitter hashtags to build amazing dynamic pages.

Let's guess that many more of the same will emerge in the months to come.

2010-03-31

societas hominum et societas rerum

Danny Ayers has recently posted a "call to arms" to try and speed up the process of adoption of semantic web technologies. And of course he has triggered the usual bunch of complaints about it. Tools are too technical, stuff is presented by geeks for geeks, data are boring, we need betteer user interfaces etc. Among many smart but technical proposals, basically adding to the general complexity issue they are supposed to solve, I will pick up this very simple one by Karl Dubost.
ACTION : Tell a story to people
I've thought about it, and here comes the best story I've come to imagine, although I'm neither a good story teller, nor good at building user interfaces. I'm just good at metaphors.
The Web is a social technology. What have been the killer Web applications so far? e-mail, blogs, Facebook, Twitter... all social stuff, whichever version of Web you call it. People understand what social entities and social links are about. So, let's tell them the story of the societas rerum (society of things) interconnected the same way as, and interconnected with, the societas hominum (society of people). Individuals connected by (meaningful) links. Yes, data are boring. Instead of the technical linked data cloud, let's show a living Web of people and things. What's in there, what it's all about : people and organisations (FOAF), places (Geonames), books (DBLP), products and services (GoodRelations), events etc. In a nutshell, the story of the Semantic Web is the story of the Social Web extended to things. And it's already there, in many ways, even if not (yet) implemented in the RDF technologies stack. Look at every web resource you get at, and ask : is this resource intended to represent and describe one definite thing? Has it a focus (see previous post)? Is it socially linked to other similar resources? If the answer is yes, then this resource participates in the societas rerum. If moreover it's linked to resources representing people, it's also participating in the societas hominum.

About the title, some will ask : why latin? Simple answer : I've been through seven years of latin classes in high school, so I have to use it somehow and show this off a little. More complex answer : Open a latin dictionary, and figure out the original scope of "societas" and "res", and if it's properly translated by "society of things". In fact I found out after forging this title that those concepts (in latin) seem to have been introduced by Antonio Gramsci. Orthodox marxists will forgive me to use them out of the original context, a bit of which is copied below. More to be found here.
One must conceive of man as a series of active relationships (a process) in which individuality, though perhaps the most important, is not, however, the only element to be taken into account. . . . The humanity which is reflected in each individuality is composed of various elements: 1. the individual; 2. other men; 3. the natural world. . . . Each one of us changes himself . . . to the extent that he changes . . . the complex relations of which he is the hub. . . . If one's own individuality is the ensemble of these relations, to create one's own personality means to acquire consciousness of them, and to modify one's own personality means to modify the ensemble of these relations. But these relations, as we have said, are not simple. Some are necessary, others are voluntary. . . . It will be said that what each individual can change is very little, considering his strength. This is true up to a point. But when the individual can associate himself with all the other individuals who want the same changes, and if the changes wanted are rational, the individual can be multiplied an impressive number of times, and can obtain a change which is far more radical than at first sight seemed possible. . . . Up to now the significance attributed to these supra-individual organisms [that the individual is related to] (both the societas hominum and the societas rerum) has been mechanistic and determinist; hence the reaction against it. It is necessary to elaborate a doctrine in which these relations are seen as active and in movement, establishing quite clearly that the source of this activity is the consciousness of the individual man who knows, wishes, admires, creates . . . and conceives of himself not as isolated but rich in the possibilities offered to him by other men and by the society of things of which he cannot help having a certain knowledge.

2010-03-25

FOAF focus

I pretty much like the new property Dan Brickley introduced as the next addition to FOAF. For what it does, see Dan's explanations yesterday on LOD forum. The rationale is clearly set:
Because conceptualisations of things as SKOS concepts are distinct from the things themselves. If this weren't the case, we couldn't have diverse treatment of common people/places/artifacts in multiple SKOS thesauri.
I let you enjoy the rest of the post, and will simply add that indeed, it's addressing in the specific context of interoperability of SKOS and FOAF the very issue we've been speaking about here for years. Things are distinct from their conceptualisations, and we need a way for various representations to focus on the same thing they are about. There is still a step further, though. The thing itself is present in the information system only through representations. The URI for the thing itself is the best proxy we can get for it in the system, but let's assume with Dan and FOAF'ers that it's somehow closer to the thing itself than concepts in thesauri, and therefore allows the latter to focus on the former. And certainly I'm delighted with the metaphore of the focus, which is yet another avatar of convergence, like the spokes of the wheel I try to keep rolling here. In french, either spokes of a wheel or converging rays of light are called rayons.

2009-11-29

Representation as translation

Words know more about themselves than we generally do, and one should always take the time to explore their etymology. Take representation, a word which is a bone of contention as soon as linguists, semioticians, ontologists, librarians and Web geeks happen to meet, and look back for a minute at its latin origin. Repraesentatio is built on praesens, which has kept in both french présent and english present its double meaning of "being here and now" and "being given". Re-presenting is therefore giving or making present again, and repetition of the praesens is important to notice. Whatever the representation stands for, signifies, means, points at, suggests ... in one word, gives or simply puts us in presence of, here and now, the praesens has been (re)present(ed) already before, somewhere else.
And since the story of representations is so old, and the beginning of the story so misty, let's assume that in the long chain of representations, each new one is leveraging previous ones, and it's turtles all the way down. So, instead of wondering about the untractable issue of the first presentation, why not focus on the process through which one representation is emerging from previous ones. This process has a name : translation.
In "Traduire - Défense et illustration du multilinguisme" (published may 2009, only available in French so far), François Ost presents the multiplicity of languages as a benediction for the life of thought, and calls for translation as our needed common language and new paradigm. This rich and thoughful book is a must read is your French is fluent enough. I cannot in one post pretend to cover all it brings about, but just re-present here a few main points, enough to show the convergence between Ost's thesis and the line of thought we've been defending here for years.
Primo, as indicated in introduction, every text or production of language is the result of a process of translation, and this process is not only happening at the borders of languages, but inside the language itself, and not only between dialects and individual expressions, but inside the mind of the locutor herself. The dialectics of the "what do you mean?" - "in other words" is not a bug of the conversation, it's definitely a constitutive and essential feature of efficient language.
Secundo, translation, at least when it's a good one, is not necessarily an entropic process through which the original meaning is degradated and betrayed. The meaning, or whatever you want to call the praesens in the original (and I will indeed call it praesens hereafter), is not necessarily lost in translation. If the translation is good, it's also present in the re-presentation. More, this new presentation is likely to enrich the original one, and maybe help to understand it better.
Tertio, to keep the praesens alive, we need to re-present it over and over again. Translation is a never-ending story. And the apparent paradox is that we need to do that the more for things considered to be untranslatable. Ost quotes there Barbara Cassin : L'intraduisible, c'est ce qu'on n'arrêtera pas de (ne pas) traduire. Saying the praesens is untranslatable means simply that no presentation is exhausting its praesens, and new re-presentations are needed to get new viewpoints enlightening the previous ones. Pretending to achieve the final representation which rules them all is falling back again into the arrogance and stupidity of Babel's Tower builders, who had lost the meaning by forgetting the diversity of languages. God's action was then not a punition, but a liberation from this deadly road of the unique thought, as Ost shows in the first chapter of the book, a brilliant hermeneutic analysis of Babel's various translations.

Porting the translation paradigm in our local metaphore is quite easy. The wheel is the never-ending story of living knowledge and languages. So many spokes, so many directions to look from, converging in the untranslatable hub which is beyond, but praesens in all re-presentations. But the translation paradigm is what was lacking to get the wheel rolling. As Umberto Eco put it about Europe : Translation is our common language. Indeed, and to put this paradigm into action, each one of us engaged in the process has to learn several languages, in order to be able to look at one language with the eyes of another. Because no one undertands her own language from inside it, and there is no meta-language to rule them all.

Last but not least, the ethical aspect of the translation paradigm : any other language is the language of one other, including mine. In translation, I learn not only to look at the other as myself, as an alter ego, which is still looking from my own viewpoint and measuring with my own metrics, but to look at myself as another, ego alter. Discovering the other-ness, the alterity inside myself, my own language, my own view of the world, is indeed a paradigm change that we all need.

2009-09-14

Asserting subclasses of open ranges or domains

I had an interesting exchange on Semantic Web list on this issue last week. You can browse the whole thread, but I would recommend answers by Pat Hayes and John Sowa.
Below are extracts of my final answer.

1. There is quite a difference to make between concepts in ontologies strongly defined by domain experts, and targeted at feeding reasoners (e.g., bio-medical or legal ontologies), and lightweight ontologies such as FOAF, VCard, Dublin Core, Geonames ... which are mainly targeting interoperability of data, and of which meaning (if not formal semantics) emerge from usage and population. I can't define formally what a Person is, but I can say that you and I are some instances.

2. For the latter said ontologies, the main objective is to provide guidelines for applications harvesting and managing data. The actual formal semantics of those models is next to nothing, but implementations can reasonably leverage them on the basis of a common sense interpretation. For example the thousands of different data models to represent a person can be re-engineered as so many specifications of the generic class foaf:Person, therefore allowing a shallow, but efficient level of data interoperability.

3. There is no more, no less semantics nor potential usability in declaring skos:Concept to be in the range of dcterms:subject, than to declare foaf:Person to be a subclass of foaf:Agent. "Dont acte"

For some other ill-defined property ranges in the Semantic Web popular ontologies, another path would be to use enumerated classes, or in a more flexible way, to indicate a published vocabulary maintaining a reference enumeration. For example when LoC publishes later this year the authoritative ISO 639-2 list of languages as a SKOS Concept Scheme, the range of dcterms:language could be restricted to the values in such a list (using e.g., a restriction on the value of skos:inScheme). This would avoid Dublin Core to go through the painful task of defining formally the class dcterms:Linguistic System which is the current specified range - with the same lack of definition as foaf:Agent. Referring to some authority is certainly the best way to deal with the issue here. We (DC) don't know what a language is, go ask ISO 639-2 folks, apparently they know because they are able to provide a list.

2009-08-10

From meaningless to ambiguous

A previous post looked at concepts as names species. Several threads about URI meaning and ambiguity later, including on www-tag@w3.org list this post from Karl Dubost, and with in mind again names and concepts as living things, here is yet another variation on morning, noon, and evening mountains.
Morning names, meaningless, useless
Names at noon, meaningful, useful
Evening names, abused, ambiguous
This basically is the life cycle of names from no use to use and eventually abuse. Use leads to meaning, then abuse leads to meaning overloading. This is the natural course of things. And since URIs are names, this is also the natural URI life cycle. So let us use what is meaningful while it is if we want meaning. Be prepared to ambiguity at the end of the day, but if evening ambiguous names have been abused, it does not mean they are useless.
Reblog this post [with Zemanta]

2009-07-30

URI species

The debate about proliferation of URIs representing the same thing keeps on rolling on various Semantic Web lists, going back again and again to the same questions. How does one discover existing URIs for a thing, if any? Is it a good or bad practice to mint a new URI for a thing which already has one? How do one link URIs identifying the same thing? Many smart and conflicting answers have been given, largely depending on the viewpoint on Web architecture and the main use of URIs in the mind of their authors. Web pragmatists and linked data evangelists tend to consider that proliferation of URIs is not necessarily a good idea, but something we are bound to live with, whereas experts in knowledge representation tend to consider it should be avoided by all means. Trust, persistence, quality of resource descriptions, use and abuse of owl:sameAs have been discussed over and over, with no obvious technical answer.

Since life provides the oldest, proven, efficient ways to store, maintain, replicate and use information, I've tried to figure if we could not learn from biology. Interestingly enough, biologists are not more able to come to a consensus about what a species is than Semantic Web gurus to agree on what is behind a URI. Somehow, the two issues are very similar. They deal with persistence of information over time. With the disclaimer that I am not a biologist, let me assume here the definition of a species as the set of individual expressions of some common genetic pool. Protection and persistence of the species genetic pool is the main occupation of any form of life. Strategies to achieve this goal present an awesome diversity, but in this variety one can find some constants. Among those are the basic facts that individuals are bound to a short life span, so the protection of the genetic pool is best achieved by assuming mortality of individuals, and ensuring duplication and replication of the information in as many individuals as possible. Not by defending a single representation behind firewalls.

How does that apply to the Semantic Web? A URI, along with the resource description it provides, can be seen as an individual expression of a species concept. As any human artefact, or any living individual, or any physical manifestation in this world, this expression is bound to be a transient. The agent who created and maintain the URI is bound to disappear, among other things. It will be less costly, as life tells us, to have copies of the information in as many expressions as possible all over the place, than to protect this specific one. Consider a URI not as the unique representation of a thing, but as an individual expression of a species.