Posted in

Big Data and the Super Bowl; Why Semantics Will Enrich Your Data

Why Include Semantics In Data? Knowledge Integration is the key. I have added a simple use case to describe the benefits. Since the Super Bowl just ended I will use a football related example.

There’s no point in adding semantics to your data if it does not provide significant benefits. One of the primary benefits of adding semantic meaning to your data is that it can be branched across domains of knowledge automatically. What do we mean with domains of knowledge?

In our example, two websites are started independently from each other. One site hosts information on current and historic Super Bowl Games; the other a large database of biographies of players that have ever played in the NFL.

Both contain complementary information in their website databases. I will cover firstly how information sharing between these sites could happen without the use of semantics. Then, describe how the same information can be shared between the two sites – and potentially beyond – with the use of semantics.

The two sites, one fronting an MS SQL database of all Super Bowl games, and another one fronting a MySQL database of all NFL players, present and past ever to appear in a Super Bowl, reside at www.superbowls.fake and www.nflpayers.fake respectively.The two sites were started independently, and do not collaborate.

The Super Bowl site lists, as its name suggests, all of the 49 games played to date and also a list of players who played in them. However, it doesn’t hold any other player information other than their name and date of birth.

The NFL players site contains a complete listing of all current and former players that have ever played in a Super Bowl, including a complete biography, plus a list of games they have played in. But, it does not contain any game details, or screenshots of the games.

Let’s look at how these two sites might collaborate under their current, more traditional data model:

  • Obviously, the users of www.superbowls.fake would benefit from being able to click on the name of a player and find out more about them – this information is stored in the MySQL database at www.nflplayers.fake.
  • Likewise, the users of www.nflplayers.fake would benefit from being able to click on the names of games that the players played in and find more information. This is stored in the MS SQL database at www.superbowls.fake.
  • Any sharing of data between the two sites cannot be done by joining tables in their databases. Firstly, they have been independently designed in the first place and so their primary keys referring to individual players or games in both databases will not be synchronized. They would have to be mapped. But secondly, they are using different database server systems which are not cross-compatible.
  • To collaborate using their current databases, the owners of either site would have to decide on a common data format by which to share information that they could both understand by using a common film and actor unique ID scheme of their own invention. They could do this, for example, by creating a secure XML endpoint on each of their websites from which they can request information from each other on demand. This way, their shared information is always up to date.

This sort of information interchange across incompatible, independently designed data systems takes time, money and human contextual interpretation of the different datasets. It also is restrictive to the data domains of only these two websites, any further additions to their knowledge from elsewhere will demand similar efforts. It requires humans to understand the meaning of the data and agree on common formats to collaborate the two databases appropriately.

With the introduction of RDF (a standard model for data interchange on the Web) and semantics, it is far easier. Let’s investigate how this could be achieved using RDF and the semantic web – it all happens automatically, not manually.
Sharing With The Semantic Web Model.

In semantic modeling, the following are important terms you should know:

  • Vocabulary – A collection of terms given a well-defined meaning that is consistent across contexts.
  • Ontology – Allows you to define contextual relationships behind a defined vocabulary. It is the cornerstone of defining a knowledge domain. A formal syntax for defining ontologies is OWL (Web Ontology Language) which is an extension to RDFS (RDF Schema), which I will explain in a follow on post.

So how do we model the two site scenario using semantic modeling? Firstly, the two sites need to apply a common, standard vocabulary to describe their data that is contextually consistent. For example, the term ‘game’ should mean the same thing for both sites, as should the term ‘player name’ and ‘player birth date’.
This may be done by the two sites adopting the same base ontology, or a common vocabulary, for expressing the meaning behind the data they expose, and publishing that data on a queryable endpoint so that the two sites can communicate with each other across the web.

With this standard vocabulary in place:

  • The two sites can now query each other using the same terms.
  • The Super Bowl game site can now query the player names on the player biographies site on-demand and gain more detail about a specific player that has played in a game.
  • The NFL Player site can now query the game details and statistics on the Super Bowl game site on-demand and gain more detail about games a player has performed in.
  • With the contextual relationships defined in a formal web ontology, further related information about the players or games, e.g. game locations, other news events happening on the same day of game or birth date or the player, or all games coached by the same coach. All may be found via the linked standard terminology without the user even imagining that information initially existed.
  • This happens without the need for transformation, mapping, or contracts being set up between the two sites. It all happens through semantics.

The good news is you often won’t have to go through the effort of defining and sharing your own ontology for your particular domain of knowledge. There are many popular, standard ontologies already distributed on the web which you can adopt, and if necessary extend yourself. I will introduce some of these in a follow on blog.

The cross-domain knowledge sharing discussed here need not just apply to websites, but also within the knowledge bases built by organizations. Semantic web technologies need not be restricted to applications or information published on the web.

Although there may be a little more time spent when first setting up a semantic database, the benefits for ease of cross-domain integration from across the globe and the time saved and ideas gained from doing so are, potentially, highly significant.

Visionary leader bringing more than 35 years of experience building successful technology business units, sales channels and companies. Have extensive experience with business startups and turnarounds, having successfully built and spun out four technology companies in the last fifteen years. Have broad experience as a CEO including heading companies in aerospace, defense contracting, telecommunications, Web 2.0 and IP video communications. Have secured funding of more than $125M for various start up companies and secured large contracts with US, European and Asian clients. During the first 20 years of my career, served in a variety of senior management positions including president of Pactel Meridian Systems, a joint venture between Nortel and Pactel.

Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.