Showing posts with label commons. Show all posts
Showing posts with label commons. Show all posts

2015-02-23

Common names, proper usage

What follows might be, as previous posts, relevant to the raging debate in and around the W3C Shapes Working Group. If you don't care too much about Latin, Greek, French, German, etymology, translation and languages at large, you can go straight to the last paragraph. But I trust my faithful readers (whoever they are) to follow me through the long preliminary linguistic meanders.

I had a while ago pointed at the enclosure of common names as trademarks. Maybe I should have written common nouns. But in French (my native language), there is a single word nom to translate both noun and name, all being cognates to Latin nomen, Greek ὄνομα, and many more avatars of the same Indo-European root. In French grammar you will say "nom commun" for "common noun" and "nom propre" for "proper noun", and a French native speaker is likely to translate in English "common name" and "proper name", both ambiguous out of context. And my purpose today is indeed to look at what it can mean for names to be common or proper beyond what it means for grammatical nouns.
Let's look into Latin again, where communis and proprius, as well as their ancient Greek equivalents κοινός and ἴδιος have roughly the semantic scope they have kept in French and English. Together they split the world into what belongs to the commons and what is proprietary or private. Beyond and before use in grammar to denote universals and particulars, further meanings have built upon good or bad characteristics associated with each term. Typically, "common" will be used as a derogatory qualifier for whatever belongs to the vulgum pecus, those common people which do not behave, think or speak properly.  The French "propre" even goes further down this derogatory path to mean "clean", with disambiguation by position ("c'est ma propre maison" = "it's my own house" vs "sa maison est propre" = "her house is clean"). Such extensions seem indeed characteristic of a language controlled by some aristocracy. It's worth noticing that the English "own" and its German cognate "Eigen" do not seem to have suffered similar semantic drifts. 
Sticking to the original meaning and forgetting the interpretations of either grammar or aristocracy, common names would be simply names belonging to the commons. Which is true, if you think about it, for just any name. A name with no community (or communality) would be useless, and actually barely a name, just a string with no shared usage and agreed-upon denotation. Under such a definition, even proper nouns are common names. From a grammatical viewpoint, "Roma" is a proper noun, but it's common to all people using it to denote the capital of Italy. To make it short, all names belong to the commons, otherwise they don't name anything at all.
The above analysis does not apply only to natural languages names (aka nouns), but also to all those technical names handled in our information system internal languages, the names used by machines to call each other in the dark (see previous post) and take actions. URIs, addresses, objects and classes names ... if those were not common names, we would have no open Web, and no open source code with reusable libraries.
But those common names, when used and interpreted by software, behave internally at run time as proper names, by all means of "proper". They each call a well defined individual object, method or whatever piece of executable code. A URI sent through the HTTP protocol is eventually calling by their internal names specific pieces of data on one or more servers, all of them running by their own, proper, often proprietary code with its idiosyncratic functional semantics.
Otherwise said, if the declarative semantics of a technical name (description of what it denotes) belongs to the commons, its performative semantics (what it does when called) is proper to the system in which it is used, and conditions at run time.

How is that relevant to the W3C Shapes debate? What this group is (maybe) seeking (or should seek) is actually a (standard) way to describe proper performative semantics for systems using RDF data. On the DC-Architecture list, +Holger Knublauch is complaining a few days ago.
Yet, there used to be a notion of a Semantic Web, in which people were able to publish ontologies together with shared semantics. On this list and also the WG it seems that this has come out of fashion, and everyone seems "obsessed" with the ability to violate the published semantics.
Violate the published semantics? Well, no, it's just about describing how the common semantics behave properly in my system. But whether that can be achieved through yet another declarative language or some interpretation of existing ones without blurring the RDF landscape a bit more, is another story. 

2013-08-09

Thou shalt not take names in vain

This is certainly too serious a subject for a Friday night in the middle of August, but that's a good time for old ideas to be written down. And indeed this has been on my mind for so long, at least since I realized that common nouns such as english timeword, windows, apple, caterpillar, shell, bull, french orange, printemps, champion, géant, carrefour, german kinder, and many more, had been "borrowed" from the language commons to become brands. This is in principle forbidden by various trademark legislations, but there are subtle workarounds. I have always considered such practices as unacceptable enclosures in the knowledge commons. They might look anecdotic, leading to rather silly cases, but some borderline practices from major Web actors show that this affair is more important that it could seem at first sight.
One could argue that the market gives back words to the commons, lists of generic or genericized trademarks are easy to find, in a variety of languages. But curiously enough,  the other way round, systematic lists of common nouns used as trademarks I could not find either in Wikipedia or anywhere else. Note sure if they could get any longer than the former, in any case the lists I proposed to start on Wikipedia were proposed for deletion a few minutes after creation by zealous wardens of the Wikipedia Holy Rules, for lack of notability of the subject. Forget about it, I'm now trying to figure how to query DBpedia to get such a list, but the distinction between a proper name and a common noun is no more explicit in DBpedia descriptions and ontology than it is in Wikipedia.
Anyway, this is not necessarily the most important aspect of the way information technologies can impact, misuse and abuse our language commons at large. There is quite a lot of rules or guidelines one could imagine for that matter, some already explicited by laws even if tricky to enforce, some yet to be specified, not to mention being enforced. There is something deeply anchored in our culture about the fair use of names, coming certainly from the way they are rooted in our religions, hence I have only a slight compunction to take inspiration below from one of the most holy and ancient set of rules. Apologies to believers who might read the following as blasphemy uttered by an old agnostic, and disclaimer to everyone else : those were not cast in stone by any god on any mountain. But if the first and main item in this list seems clearly inspired by the Third Commandment, well, yes it is, and not only in form. The underlying claim is that every word, every name, carries along with it enough history and legacy to be honoured. Those who don't care that much about such religious considerations can read this as pragmatic deontological guidelines for a fair, efficient and sustainable use of names in our information systems at large, and on the Web in particular.

Here goes, ten items of course to stick to the original format. 
  1. Thou shalt not take names in vain
  2. Honour the many meanings of a name, for they belong to the Commons
  3. Acknowledge linguistic and semantic diversity, polysemy and synonymy
  4. Do not steal names from the Commons to be your proper names
  5. Do not sell and buy names, for they belong to the Commons
  6. Do not hide yourself or your products under false names
  7. Do not use names against their common meaning 
  8. Do not enforce your own meanings upon others
  9. Expose your meanings to the Commons, for they will be welcome
  10. Share your own names with the Commons, for they will thrive forever
I won't dwelve today in the details of each of those, some might look quite cryptic and need to be expanded in further posts. Just a remark on the first (and most important) one. The "take in vain" used by the King James version of the Bible has been replaced in more recent translations by "misuse". I prefer the former, which conveys the notion that whenever you use the name, it's not for nothing or something without importance and consequence. When you use a name, you should have well thought about its meaning. In French you would translate at best "Tu ne prendras pas les noms à la légère."

2013-01-08

Everyone knows what a Semantic Dog looks like

A fierce debate is raging those days on the lod-public list between +Kingsley Idehen and +Hugh Glaser, plus a couple of others. The question is to know if the huge efforts to publish mountains of linked data have produced so far any kind of visible and useful applications consuming them. In other words, where are the semantic dogs consuming those heaps of Semantic Web Dog Food? Kingsley holds it of course that they have invaded all the Web avenues, and Hugh that they are nowhere to be seen. Obviously they are not looking for the same kind of dogs, or they don't agree on what a semantic dog could look like, making me wondering if such dogs might be akin to the dragon of the story.
This debate is to compare with +Amit Sheth's recent post entitled "Data Semantics and Semantic Web - 2012 year-end reflections and prognosis" suggesting among other things that although another five years or so could be necessary for the Linked Open Data to gain enough quality to allow building upon it seriously, on the other hand things like Google Knowledge Graph are opening the way to pervasive semantic applications. 
Beyond those ongoing cries of  "Publish more linked data" and "Show me the applications" why not try in this New Year time to think about linked data in a Long Now perspective?

2012-07-10

LOV moves to OKFN


The LOV story enters a new area today. Pierre-Yves and myself are pleased to announce that the Linked Open Vocabularies (LOV) project is now hosted at OKFN. Many thanks to the Open Knowledge Foundation team which provides the technical support which will help to ensure a sustainable future for this project and the Vocabulary Commons at large.
Pierre-Yves has done a tremendous technical work, not only to ensure a smooth migration from former Mondeca Labs hosting, but also to add a new vocabulary versions and history feature. See for example Dublin Core or SKOS. Versions are displayed on a timeline and the version files and change notes are available as long as we could put our hands on them. Thanks also to all creators and publishers who helped us to sort out the history of their vocabularies. This is far from done for all vocabularies, but we hope the examples already there will push other publishers to make this history available. 

2012-04-06

LOV stories, Part 3: Vocabularies as Heritage

In previous posts we introduced the Vocabulary Commons, following some of its Gardeners. Let's now  try to imagine how they can turn into a sustainable and resilient ecosystem. In the current state of affairs they look more like a young forest pionneering a new land. A lot of opportunist species have invaded the landscape, some are conspicuous and seem here to stay, some look like they have no future, others have set up in small niches, all kind of interactions and dependencies have emerged.
Should we let the invisible hand of natural selection operate, and let the fittest survive? Do we want those commons to become a wild messy jungle, or rather a pleasant, useful and sustainable garden that we, our children and their children will enjoy? It's certainly how to achieve the latter the DCMI Vocabulary Management Community group has in mind. This group will hold its kick-off meeting on the first day of the London Seminar : Five Years On at the end of this month.

2012-03-22

LOV stories, Part 2 : Gardeners and Gatekeepers

In the previous post, we have seen why Linked Open Vocabularies should be managed following the principles and rules of the commons : shared resources in which each stakeholder has an equal interest
Which means that in theory at least, all stakeholders of the commons should be both users and gardeners. In the vocabularies commons the stakeholders (and hence potential gardeners) are as various as can be, actually they encompass all the actual or potential providers, curators or users of linked data, since there is no quality linked data without quality linked vocabularies, as we have explained in a recent post. But let's look at who are the actual current gardeners of the vocabulary commons.

2012-03-15

LOV stories, Part 1 : The Commons

It's now been about one year of work with my colleague Pierre-Yves Vandenbussche on the Linked Open Vocabularies (LOV) project. I've already mentioned it lately here on this blog and in various conversations on Google+. This is the first on a series of posts where I will elaborate a little more about the general vision and philosophy of this project, lessons learned so far, and explore possible roadmap towards its sustainable future. 
The philosophy of LOV in a nutshell is the Philosophy of the Commons.

2011-12-27

A Web of unfinished weavings

The Web is full of enthusiastic beginnings. Regular and steady follow-up, such as Astronomy Picture of The Day, of which daily archives are available since 1995, are harder to find. The statistics of this blog, and many more of the same, provide typical examples, but unfinished weaving is unfortunately not limited to personal looms, it's also undermining greater collective endeavors. I was looking today at the state of some Wikipedia articles I'd been seriously contributing to, five years ago, such as the one about  SKOS, and figured they need to be seriously updated. But if it's a lot of fun starting a new article, it's quite a boring task to go through it five years after, cleaning and updating it. And since it's a collaborative task, someone else could care after all. Many interpretations have been given to the fact that many people have given up editing Wikipedia, such as growing complexity, bureaucracy, edit wars etc. But people can cope with all this, as long as there is fun, and as long as there is something new every day. Wikipedia is now more than ten years old. It started in the previous century. At Web time scale, it's a very old-fashioned thing.
I see a similar trend undermining the Semantic Web. One could think that vocabularies used by linked data, since more and more people and application rely upon them, would be maintained and curated like precious assets. Actually after a year or so of exploration of this ecosystem, trying to federate the community around its crucial importance, I'm surprised that many of those vocabularies just sit there on a Web shelf, letting to everyone's guess if they're here to stay, if they have been or will be updated, if their publishers have any roadmap for their future evolution, or even if they still remember them.
Weaving the Web? If the weavers seem to be attracted every day by the next trendy loom, and forget to finish what's up on the old ones, the tapestry will always look like an unfinished patchwork. Is this the knowledge we want to build?

2006-06-20

Wikipedia's semantic cow paths

I've been quite silent here on this blog for two months, and meanwhile resumed a bit of my Wikipedian activity lately. Although I'd been an enthusiastic early adopter of Wikipedia back in 2001, I have a poor and episodic editing history so far. But every time I've been coming back to the editor dashboard after months or years of inactivity, I've been amazed by the tremendous qualitative growth of the toolkit made available to users. Wikipedia's growth has been stressed again and again in terms of quantity and quality of articles, languages, editors, popularity as a reference resource etc. But was has not been stressed enough is the parallel growth in terms of features supporting better search and editing. And many of those features are in fact adding a quality of information which makes it ready for semantic parsing, and easy RDF re-writing. The basis of it all is a sound use of URIs, names and namespace.
For example http://en.wikipedia.org/wiki/Volcano defines without ambiguity the unique page dedicated to this geological feature, while http://en.wikipedia.org/wiki/Volcano_(disambiguation) is a hub to potential homonyms such as http://en.wikipedia.org/wiki/Volcano_(film). Links are provided to similar resources in other languages such as http://fr.wikipedia.org/wiki/Volcan. One can reasonably use any of those URIs to identify the concept "Volcano". But what about a class Volcano? Here come Wikipedia categories. http://en.wikipedia.org/wiki/Category:Active_volcanoes is ready-made for a class, and parsing the page will give you easily the list of instances, with a link to the full description page. This description itself is formatted using templates such as "infoboxes", so that a page in a category "Active Volcanoes" will yield standard properties in a standard format. Easy to turn this into a data base, and if one interprets the infobox elements as so many properties, turn it into a RDF description, and the infobox structure itself in a RDFS or OWL description of the matching class.
What should we learn from that? That from a collaborative and mostly non-directed process, are emerging cow paths which look more and more like semantic markup, ready to be spidered by smart parsers and tools, either to improve Wikipedia content itself (there are already a bunch of bots and agents doing that), or to extract of this amazing knowledge base any kind of structured data in whatever format, including implicit ontology, like structure of categories, attributes used, and the like. And my hunch is that this process will be quickly much more effective for Semantic Web building than many costly academic ontologies nobody will ever use.

[2010-04-08] : What is described here is exactly what DBpedia started in 2007

2005-02-03

Wikipedia URLs a Subject Codes

The link is to David Megginson's blog. This struck me as terribly interesting.
Over in my aviation weblog, I find myself more and more linking to Wikipedia whenever I’m discussing a concept, person, place, or anything else that doesn’t have its own, canonical home page. If, as I suspect, lots of other bloggers are doing the same, then links to Wikipedia articles may soon be the blogsphere’s answer to subject codes.
The idea follows something similar from James Tauber, who points to a tagging scheme from Technorati. In the end, it seems that subject identity lies in the realm of concensus or agreement; Wikipedia appears popular enough that its URLs might serve at least one important aspect of the subject identity issue.