Showing posts with label language. Show all posts
Showing posts with label language. Show all posts

2015-04-11

From names to sentences, the Web language story.

Conversation about text and names and how they are interwoven within the Web architexture is going on here and there. The more it goes, the more I feel we need more non technical narratives and metaphors to have people get what the (Semantic) Web is all about. We have drowned them under technical talks and schemas of layers of architecture and protocols and data structures and ontologies and applications ... and the neat result is that too many of them, and smart people, think only experts, engineers and geeks can grok it. So let me try one of such - hopefully simple - narratives. 

The story of the Web is just the story of language, continued by other means. Forging names to call things, and weaving those names in sentences and texts. On the Web, things have those weird names called URIs, but names all the same. As we have seen in a previous post, a name is to begin with a way to shout and identify people and things in the night. On the Web to call a thing by its URI-name you will use some interface, a browser, a service, an application, and at this call something will come through the interface. Well, the thing you have called does not actually come itself to you through the network, but you get something which is hopefully a good enough representation of the thing. The deep ontological question of the relationship between the name and what is named has been discussed for ages and will continue forever. The Web does not change that issue, does not solve it, just provides new use cases and occasions to wonder about it. But this is not my point today.

On the first ages of the Web, calling things was all you could do with those URI-names. You had the language ability of a two years old kid. You could say "orange" or "milk" when you were thirsty, and "dog" and "cat" and "car" and "sea" and "plane" when you saw or wanted one, and cry for everything else you could not express or the dumb Web would not understand. With no more sophisticated language constructs, you could nevertheless discover the wealth of the Web, through iterative serendipitous calls. Because courtesy of the Web is such that when you call for a thing the answer comes often back with a bunch of other names you can further call (an hyperlink just does that, enabling you to call another name just by a click). You would bring back home things you had not the faintest idea of the very existence a minute before. Remember this jubilation, the magic of your first Web navigation, twenty years ago? Like a kid laughing aloud when discovering the tremendous power of names to call things.
Today in many (most) of our interactions with the Web we are no more aware of using names. We make actions with our fingertips, barely guessing that under the hood, this is transformed in a client calling a server or something on this server by some name, and many calls are made on the network to bring back what your fingers asked. Only geeks and engineers know that. The youngest generations who have not known the first ages of the Web, and interact only through such interfaces, are plainly ignoring all that names affair. Did you say URL Dad? What's that? It sounds so 90's ...

Now when you grow older than two, you go beyond using names just for shouting them in the face of the world, you begin to understand and build yourself sentences. That's a complete new experience, a new dimension of language unfolding. You link names together, you discover the texture, the power to understand and invent stories and to ask and answer questions. You still use the same names, you are still interested in oranges, cats, dogs and cars, and all the thousands of things which are the children of naming. But you are now able to weave them together using verbs (predicates), qualifiers and quantifiers and logical coordination. You have become a language weaver.

And that's exactly and simply what the Semantic Web is about, and how it extends the previous Web. Just growing and learning to weave sentences, telling stories, asking questions. But using the same URI-names as before. Any URI-name of the good old Web can become a part of the Semantic Web. Just write a sentence, publish a triple using it as subject or object, and here you are. 

2015-03-02

Could computers invent language?

Artificial intelligence is something about which not a line has been written in these pages in next to two hundred posts and over more than ten years. But I feel today like I should drop a couple of thoughts about it, after exchanges on Google+ around this post by +Gideon Rosenblatt and that one by +Bill Slawski, not to mention recent fears expressed by more famous people.
There are many definitions of artificial intelligence, and I will not quote or choose any. Samely, popular issues I also prefer to let alone, like knowing if computers are able to deal only with data and algorithms, or if they can produce information or even knowledge, or if they think and can individually or collectively accede to consciousness or even wisdom. All those terms are fuzzy enough to allow anyone to write anything and its contrary on such issues. Let's rather look at some concrete applications.
Pattern recognition is one of the great and most popular achievements of artificial intelligence. Programs are now able with quite good performance to translate speech into written language, identify music tracks, cluster similar news, identify people and cats on photographs etc. 
Automatic translation is also quite popular, and working not that bad for simple factual texts, has still hard time dealing with context to solve ambiguity, understand puns and implicit references, all things generally associated with intelligent understanding of a text. 
Question-answering is also making great progress, based on more and more rich and complex knowledge graphs, and translation of natural language question into formal queries.
No doubt algorithms will continue to improve in those domains, with many useful applications and some related and important issues regarding privacy and delegation of decision to algorithms.

All the above tasks deal more or less with the ability of computers to process successfully our languages. But, and this is where I'm bound from the start, there is a fundamental capacity of human intelligence which, as far as I know, has not even began to be mimicked by algorithms. It's the capacity to invent language. It has been largely discussed since Wittgenstein whether a private language is possible or not, but there is no discussion that language has been and still is built collectively through a proceess of collective continuous invention. Anyone can invent a new word or a new linguistic form; whether it will be integrated into the language commons depends of many criteria akin to the ones enabling a new species to expand and survive or disappear. This is the way our languages constantly evolve and adapt to the changing world of our communication and discourse needs. Could computers be able to mimick such a process, take part in it, and even expand it further than humans? Could algorithms be able to produce new and relevant words, smoothly integrated in the existing language, to name concepts not yet discovered or named? In short, are computers able to take part in the continuous invention of language, and not only make a smart use of the existing one?
Such a perspective would be indeed fascinating and certainly scary, insofar as machines inventing collectively such language extensions would not necessarily share them with humans, and even if they do, humans would not necessarily be able to understand them

Whether such an evolution is possible at all or in a foreseeable future is a good question. Whether we should hope for it and work to let it happen, or should fear and prevent it, is yet a more interesting one. But at the very least, those questions we can technically specify, making them much more valuable for assessment and definition of artificial intelligence than vague digressions on whether computers can think, have knowledge or can become conscious. We don't even really know what the latter means for humans, our shared language being the closest proxy we have for whatever is going on in our brainware. So let's assess the progress of artificial intelligence by the same criteria we generally use to assess the human intelligence, its ability to deal with language, from plain naming of things to invention of new concepts.

2014-12-30

Untranslatables of philosophical engineering

It's been five years now since I first wrote here about translation and untranslatables. This paradigm has been on my mind ever since. Thanks to Santa Claus, the Dictionary of Untranslatables - English edition ten years after the original publication in French - has found its way to my bookshelves. This is a huge reference collaborative work which I could not recommend enough to lovers of languages and philosophy. That includes hopefully about anyone reading those lines.
By design, this work is a never-ending story, the untranslatable being, as Barbara Cassin keeps reminding us, what needs to be translated again and again, and we'll never be done with it. That's why adaptation of the Dictionary in more languages is planned or in the making.
It's interesting to explore the Dictionary to see if and how it helps us to understand the terms of our philosophical engineering realm, as Alexandre Monnin  and  Harry Halpin among others now like to call it. Those folks have already done a great job trying to shed a bit of philosophical light over our passionate permathreads about identity, reference, resourcedocument, work, concept and more of the same. If we are never done with those, it's not necessarily because philosophical engineers are either lousy philosophers or crappy engineers, or even both. Granted, all sorts of people have been involved in such debates, often with great innocence and not enough either philosophical or technical background to tackle such difficult issues. But many more are very good philosophers or smart engineers indeed, and a good bunch of them can proudly claim to be both. If those smart brains can't come after years of debate with definitive and clear agreements about such concepts, and how they should be translated in the Web languages and implemented in its very architecture, it's certainly because they are akin to the hard core of untranslatables that philosophers and linguists have kept trying to translate for ages. And with no surprise, some of them are already important entries in the said Dictionary, because they have a long history in pre-Web philosophy, like identity, reference, thing, work, representation, or word. Some more have minor or side entries, like topic or description, and some are absent because their conceptual difficulty has emerged as a pure product of the Web, like resource, content, or data
Translation issues for such hard-core concepts do not only happen when translating them into other natural languages than their original one, most of the time some dialect of Globish. They also happen any time one wants to cross-link or make interoperable different business, technical or domain dialects, even using the same (apparently) natural language. Classes of librarians, knowledge engineers bred in Description Logics and object-oriented developers are different animals, their actual operational semantics are radically different, but close enough to bring about potential confusion when those people try to sit and build something together, an event likely to happen in the context of the Web. The local usages all appear as avatars of some (the same, or not) fuzzy underlying untranslatable, having to do with hierarchy and inheritance, rooted in the pervasive genus-species paradigm, and bearing the same name here and there by chance and necessity of linguistic evolution, involving various contingent reasons such as finitude of lexicons, laziness or reluctance to coin new and specific terms, jokes and puns, lousy semantic extensions etc. 
If the Web is here to stay as a major production of human knowledge, and considered as such by philosophers, a future revision or extension of the Dictionary will (should) include those untranslatables of philosophical engineering. For the greatest benefit of both philosophers and engineers, and the delight of all those pretending to be both.

2009-11-29

Representation as translation

Words know more about themselves than we generally do, and one should always take the time to explore their etymology. Take representation, a word which is a bone of contention as soon as linguists, semioticians, ontologists, librarians and Web geeks happen to meet, and look back for a minute at its latin origin. Repraesentatio is built on praesens, which has kept in both french présent and english present its double meaning of "being here and now" and "being given". Re-presenting is therefore giving or making present again, and repetition of the praesens is important to notice. Whatever the representation stands for, signifies, means, points at, suggests ... in one word, gives or simply puts us in presence of, here and now, the praesens has been (re)present(ed) already before, somewhere else.
And since the story of representations is so old, and the beginning of the story so misty, let's assume that in the long chain of representations, each new one is leveraging previous ones, and it's turtles all the way down. So, instead of wondering about the untractable issue of the first presentation, why not focus on the process through which one representation is emerging from previous ones. This process has a name : translation.
In "Traduire - Défense et illustration du multilinguisme" (published may 2009, only available in French so far), François Ost presents the multiplicity of languages as a benediction for the life of thought, and calls for translation as our needed common language and new paradigm. This rich and thoughful book is a must read is your French is fluent enough. I cannot in one post pretend to cover all it brings about, but just re-present here a few main points, enough to show the convergence between Ost's thesis and the line of thought we've been defending here for years.
Primo, as indicated in introduction, every text or production of language is the result of a process of translation, and this process is not only happening at the borders of languages, but inside the language itself, and not only between dialects and individual expressions, but inside the mind of the locutor herself. The dialectics of the "what do you mean?" - "in other words" is not a bug of the conversation, it's definitely a constitutive and essential feature of efficient language.
Secundo, translation, at least when it's a good one, is not necessarily an entropic process through which the original meaning is degradated and betrayed. The meaning, or whatever you want to call the praesens in the original (and I will indeed call it praesens hereafter), is not necessarily lost in translation. If the translation is good, it's also present in the re-presentation. More, this new presentation is likely to enrich the original one, and maybe help to understand it better.
Tertio, to keep the praesens alive, we need to re-present it over and over again. Translation is a never-ending story. And the apparent paradox is that we need to do that the more for things considered to be untranslatable. Ost quotes there Barbara Cassin : L'intraduisible, c'est ce qu'on n'arrêtera pas de (ne pas) traduire. Saying the praesens is untranslatable means simply that no presentation is exhausting its praesens, and new re-presentations are needed to get new viewpoints enlightening the previous ones. Pretending to achieve the final representation which rules them all is falling back again into the arrogance and stupidity of Babel's Tower builders, who had lost the meaning by forgetting the diversity of languages. God's action was then not a punition, but a liberation from this deadly road of the unique thought, as Ost shows in the first chapter of the book, a brilliant hermeneutic analysis of Babel's various translations.

Porting the translation paradigm in our local metaphore is quite easy. The wheel is the never-ending story of living knowledge and languages. So many spokes, so many directions to look from, converging in the untranslatable hub which is beyond, but praesens in all re-presentations. But the translation paradigm is what was lacking to get the wheel rolling. As Umberto Eco put it about Europe : Translation is our common language. Indeed, and to put this paradigm into action, each one of us engaged in the process has to learn several languages, in order to be able to look at one language with the eyes of another. Because no one undertands her own language from inside it, and there is no meta-language to rule them all.

Last but not least, the ethical aspect of the translation paradigm : any other language is the language of one other, including mine. In translation, I learn not only to look at the other as myself, as an alter ego, which is still looking from my own viewpoint and measuring with my own metrics, but to look at myself as another, ego alter. Discovering the other-ness, the alterity inside myself, my own language, my own view of the world, is indeed a paradigm change that we all need.

2007-10-09

URIs for languages

I've eventually given up trying representing hubjects at all, at least for the moment. I had a serious try at it at lingvoj.org. But after discussions in the Linking Open Data forum, I eventually surrendered and published the languages description in a way conformant to W3C recommandations for Semantic Web architecture, with content negociation, 303 redirects and the like. I've even suppressed the previous post here saying otherwise, which would be now full of dead links and would bring about confusion.
So we'll see how this flies. Feed your favourite tool with the URI http://www.lingvoj.org/lang/zh, and figure by yourself if it provides a useful description of the Chinese language, both for humans and machines.

2007-04-13

A bit of Chinese

I've been dreaming about learning a bit more of Chinese than the few characters I've been playing with ever since I discovered the excellent introduction by Kyril Ryjik's "L'Idiot Chinois", around 1980 or so (this book is unfortunately sold out now). I've decided that now is the time to learn Chinese, for all sorts of obvious reasons. Hence the extract of the original Chinese text of the now famous "wheel and hub" quote in this page. If your browser does not support Chinese characters, you should do something about it right away.
I've tried to come out with my own translation. Note that I got rid of any capitalization, because Chinese has no notion of Absolute Capital Things. The concepts represented by Chinese characters are most of the time both concrete and abstract, generic and specific. No definite or indefinite articles, no clear distinction between grammatical nature and function. Depending on the context, the same character may be translated as a noun, a verb, or adjective.
In a way, Chinese characters are by nature ... hubjects.

2006-12-27

A couple of things I've been about lately

I've been silent here for over two months now, my blogging time devoted to the Mondeca blog in French Leçons de Choses. But there is a couple of things I've been working on, worth mentioning.

I've exchanged with Michel Biezunski on his Data Projection Model , and found out that its genericity and simplicity made it easy and straightforward to express the structure of Mondeca ITM, without the borderline hacking needed when using either OWL-RDF or XTM for the same task. Now open questions: What will happen with that model? Who will see the benefits over languages already in this space, and singularly over RDF? Who will build tools supporting it?

Been wondering if a semiotic approach could shed some light on our thoughts on referents, and came out with a RDF semiotic triangle. The URI is the signifier, the RDF description is the formalisation of the signified concept associated with the URI. The referent is out of the language and signs realm, and should stay there. In this approach, attempting to achieve a representation of the referent, even using tricks as blank resources or hubjects of any kind, is therefore a recursive trap and actually a non-sense. So any declaration of same-ness or identity of referents should be avoided. Only concepts bear identity, not their referents. From that point on, came to the idea that linking different concepts/signs (URI + RDF description) which humans consider to have more or less similar referent will take the form of processing rules, more than declarative semantics.

Thanks to Jakob Voss for this post in a long thread on public-esw-thes list, which really triggered a kind of illumination about this. As an example, trying to say that my SKOS concept a:Restaurant has the same referent as your OWL class b:Restaurant through any RDF declarative relation between those two resources shoud be avoided. But I can set in my system a functional rule expressing that any document of which subject is an instance of your b:Restaurant class will be indexed against my a:Restaurant concept. The referent is represented nowhere, but it is acting at the core of this rule.

Actually we have this very indexing rule mechanism working in some Mondeca applications, and I have submitted a paper to XTech 2007 about it. More to come if ever the paper is selected.

Lately, got interested again in triggering some process to have languages available not only as tags to use in XML, but as proper RDF resources. This is an old story tracking back to OASIS Published Subjects Technical Committees, and singularly PSI for languages. Track this topic on ESW Wiki, and see here for ongoing thread and more explanations. There again, my proposal is to forget absolute identification of a language by a URI. Concepts identified by URI are the properties and property values than can be declared for a language, and let applications decide on which properties are useful to them. No absolute rule saying that two descriptions refer to the same language.

2005-10-10

The Search for the Perfect Language

Started diving in this fascinating book by Umberto Eco last week-end. Discovered the French translation available in my local library. Really worth reading, to understand that what we are doing here and in many other places today is just another episode of a very long story. A quote among many, this one from Descartes in a letter to Marin Mersenne in 1629 (my own translation from French, hope it makes sense).
I take that such a language is possible, and that the science on which it depends can be found, by mean of which farmers could best grasp the truth of things than philosophers do today. But don't hope to ever see it in use; that would suppose great changes in the order of things, and would need the world to be a heaven on earth, something worth to propose only in the world of novels.

2005-07-29

Fun with Hubjects

What happens when you are standing in the shower at oh-dark-thirty in the morning and you start thinking about hubjects? It goes like this.

A hubject is the result a phonetic accident when two memes, subject and hub, have a translocation error performed on them. This accident is part of what we now call directed evolution. By contrast, the philadelphia chromosome, known to be behind several cancers, the most prominent being chronic myelogenous leukemia, is a translocation error, not thought to be directed, between chromosomes 9 and 22. That error splices part of 9 with part of 22 into the famous BCR-ABL splice, characteristically referred to as "ph+" (because it was the first cancer gene discovered -- in Philadelphia -- following Watson&Crick). But, there remains the other parts which did not become famous, but which also get together. A dissertation at UCLA showed that object to be benign.

So, what's that got to do with hubjects? That's what hits when you've got soap in your eyes. In America, we have this whole thing about suburbs. Stay with me here; don't try to guess where this is going. We speak of living in the 'burbs. Could we say that a SubjectProxy (aka: Topic) living in a topic map is, um..., living in the 'bubs? Only if we had a different phonetic accident on the same memes and came up with, brace yourself, sububs.

You know, we can do that. Language is the longest running open source project in the entire universe. We can make up names for things till the cows come home, and beyond. At the core, however, the identity of the subject remains the same. Go figure.

2005-05-23

John Udell; delicious; language evolution

Click on the title and you'll hear and watch John Udell walk you through using "delicious". It's a lucid exposition of how tags and tagging contribute to language evolution. Tagging, as Bernard suggests in his previous univers immedia post here, appears to have great merit. John Udell's page ends with a thoughtful commentary on the relationship between tagging and Steven Pinker's writings on language.

Consider this: language is the longest-running open source project on this planet.

2004-10-29

Vocabulary Management for the Semantic Web

Tom Baker, coordinator of the Vocabulary Management Task Force in the W3C SWBPD WG has released a first draft of the technical note which is the intended deliverable of the VM TF, shortly defined in the TF definition page as following:
A relatively concise technical note summarizing principles of good practice, with pointers to examples, about the identification of terms and term sets with URIs, related policies and etiquette, and expectations regarding documentation.
The current draft structure brings interestingly together practices from various communities and languages : FOAF, Dublin Core, SKOS, WordNet ... and Published Subjects. The release also includes an exhaustive list of links on the subject.The WG and the TF will have their F2F meeting next week in Bristol. I've been put under the interesting challenge to deliver something for the following section:


Section 3.3.
What does it mean to "use" Terms from one Vocabulary in another?
The problem of "semantic context". Terms may be embedded in clusters of relations from which they may be seen in part to derive their meaning. It may therefore not always be sensible to use those terms out of context. Examples include the terms of thesauri or ontologies, as well as XML elements, which may be defined with respect to parent elements and may therefore not always be reusable as properties in an RDF sense without violating their semantic intent.
In other words : does a term, clearly identified by its URI, but used in a context different from its original one, identify the same original concept?

2004-10-20

Typing : the main context in identification process?

Looking back at some recent posts, seems that most of the time, identification processes take place in contexts where the type (class, category ...) of the thing to identify has been explicitly or implicitly defined. It's explicitly said in the paper quoted in my yesterday post, and it appears in an ongoing thread on topicmapmail, where Lars Marius Garshol, speaking about Subject Identity Measure, writes:
"Another consideration is that I think types are extremely important. If the names are the same but the types are disjoint (person and place, say) then you can safely ignore the names. You might even want to make the algorithm consider typing topics first, and only afterwards go after the instances."
This is also maybe behind Jack's post "Seeing As..." . You can identify a "real" person only if you see it as a person, and on the other hand you can consider, given the appropriate identification context, any kind of data : a photograph, a phone number, an email address, a handwriting, the sound of a voice, a perfume ... as identifying a person if you see those data as persons, meaning that you have set an identification context where the type of object to identify is "Person".

Those considerations make me wonder about the credibility of URIs, PSIs and other kinds of "universal identifiers", if set outside any processing context, and maybe the minimal processing context should be the type of the thing identified. If http://psi.oasis-open.org/iso/639/#fra is set as a PSI for an instance of the class "Language", would it make any sense to use this identifier to identify a topic, without assuming implicitly that this topic is itself an instance of this same class?