Welcome!

Open Source Authors: Carmen Gonzalez, Sandi Mappic, JP Morgenthal, Phil Whelan, Rex Morrow, Datical

Related Topics: Linux, Open Source

Linux: Article

Linux Cover Story — Multidimensional Tagging

Finding information naturally

Multidimensional tagging, a key component in social sharing sites, can potentially help enterprises manage large stores of information. In this article, I'll examine the ways that multidimensional tagging will be implemented using Open Source tools.

As storage costs have continued to decrease, organizations create and retain more information. It is easy - and useful - to keep information online, just in case it's needed. The challenge has become managing information not simply storing it. Structured systems, such as relational database management systems (RDBMS), have well-developed tools for indexing, locating, and retrieving information. Unfortunately not all information fits well into structured systems. Graphics, presentations, spreadsheets, and word processor documents have all proliferated, no longer bound by storage costs. This is especially true of desktop systems where 200GB hard drives are common.

Finding exactly the information needed at the right moment has become difficult. A typical desktop drive has tens of thousands of files. A file server or NAS device could easily have hundreds of thousands or even millions of files. Even if system files and applications are discounted, that's still a lot of files to wade through.

File Systems and Search Engines Provide Some Relief
File systems impose a certain degree of order on unstructured information. By allowing information to be placed in hierarchical directories, file systems group files that share some relationship into meaningful categories. The downside to categorization using a file system is complexity for the user. The user is forced to descend through more and more layers of directories to find what he wants. Just remembering where the correct directory is becomes a chore. Information organization via the file system becomes less useful when information is spread across an enterprise. Incompatible systems, individual ways of building directory structures, and shear scale make this a difficult way to manage information assets across an organization, even a small one.

Search engines provide some respite from our information organizational woes. By indexing keywords and storing references to their source in a database, it's possible to look quickly through all the available files. The vendors of searching engine technology, drawing from vast experience in indexing hundreds of millions of Web sites, provide tools that let users find information spread across a huge number of desktop disks and enterprise storage systems.

Tagging Provides Necessary Clues
The trouble with the typical search engine is that the user must first know what he's looking for. He has to have some idea what keywords are indexed for a particular piece of information. If the indexed words don't match the words that a user thinks apply to the information, then the search engine won't find what the user is looking for. A marketing brochure, for example, may not have any words in it that say "marketing brochure" but that's how the user thinks of it. The way the information is categorized in the users head doesn't always match the strict keyword index of the engine.

This has already become something of a problem for Web sites that allow users to share large amounts of information. Cues are needed to help visitors who come to a site find what interests them without knowing the exact nature of the content. In typical Internet fashion an organic solution has arisen called multidimensional tagging, labeling, social tagging, folksonomy, or just simply tagging. It's multidimensional because a single piece of information can have many different tags, reflecting the different dimensions that users apply to it. Tagging lets users assign a set of categories to a piece of information when they create it. The tag system, also known as the tag cloud, grows as people use the information and see relationships in it that the original author might have missed. Users categorize information according to how they view the information, which makes it useful for groups of people who don't always think alike, such as engineers and marketing people.

Tagging is a key feature in social sharing sites such as Yahoo's del.icio.us and Flickr as well as You Tube, and Userscripts.org. Whether it's sharing interesting Web pages, photos, video, or Greasemonkey scripts, all of these sites rely on user categorization. Without tagging, no one would find anything of interest and the site would fail. Unlike simple storage sites (such as Yahoo Photos and Yahoo Briefcase), they require a way of presenting information to users that lets them find it quickly. Users can find information even when they're not really sure what they want. Social sites don't abandon search engines. Instead they integrate searching with tagging to provide a breadth of information retrieval options. Most let users search the tag cloud as well as scanned keywords, providing a rich search environment.

Tagging: New to the Enterprise
Tagging technology for the enterprise environment is new and not widely deployed in products. That's unfortunate. Not only is it extremely useful for finding information, it's also a natural way to do it. This is especially true for people used to social sharing sites. The tagging methodology facilitates the efficient sharing of information across many users and a large file space. It's exactly what enterprises need to make best use of the great stores of unstructured information on their corporate networks.

There are three ways that tagging is being implemented in corporate environments: integrated into applications; as a part of a standalone information management system; and, eventually, as a file system feature. The first kind of implementation is readily available. Image management systems, even ones directed at the desktop environment like Google's Picasa, include tagging as a core feature. The next version of the Thunderbird e-mail client (version 2.0) is expected to include e-mail tagging, augmenting its current search capabilities. Of course, once users get used to tagging for managing certain types of information, they will wonder why they can't use it for all the information that they need to access. They'll expect tag clouds that span all kinds of information in the enterprise.

Tagging is also being implemented in targeted information management tools. Tools for searching large stores of information in a corporate network are still at an early stage but tagging should be expected to make an appearance in information management and search engine tools in the near future. Consider this, Yahoo uses tagging in Flickr and del.icio.us as well as the upcoming My Web 2.0. Is it a stretch to expect it to implement tagging in the corporate search arena? The same is true for Google, which uses tagging on its eBlogger site, GMail Web-based mail service, and Picasa image management tool.

Finally, tagging can be expected to become a feature of the file system and operating system. Some aspects of tagging already exist in operating systems. The ability to attach keywords to files in Microsoft Windows is an example of file system-level tagging. These keywords are currently read by Microsoft's desktop search engine, creating a crude multidimensional tagging feature. Of course, entering and displaying tags is clunky and tags can't be displayed in and of themselves, rendering it more of a hack than a real feature. It does, however, point the way to future features of the operating system. Fully integrated into an operating system as normal metadata and using standard visual cues such those used on social sites, tagging will become a typical part of most corporate environments.

Tools Exist, File System Hooks Don't
The tools for corporate tagging capabilities already exist in the Open Source community. Most of it is encapsulated in the tools used by social bookmarking sites, which are often based on the LAMP stack. They're typically written in common scripting languages, such as Perl or Python, or Java. One such Open Source tool is unalog. Ostensibly a social bookmarking system, it's written in Python and the source is readily available on SourceForge. While the core tools exist, the hooks into the file system are still mostly missing.

A somewhat different but innovative approach is evident with Flickrfs or the Flickr File System. Based on FUSE, it creates a virtual file system with tagging for the Flickr digital photo management service. A fusion of file system and service, Flickrfs lets Linux users access the Flickr service as if it were any other mounted Linux file system. Photos can be accessed through the same tags available on Flickr using standard Linux commands such as cp. Flickrfs represents another way that tagging may come to information management - as a specific application or service but integrated into the normal file system.

Conclusion
Multidimensional tagging provides an opportunity to let users manage information more in line with their natural way of thinking. By sharing tags across the enterprise, users will spend less time looking for information and more time making use of it. Unlike other collaborative systems, users do all the work without legions of editors making decisions that users find mystifying. The social sites on the Internet have shown this to be a viable information management model. It's a matter of how and when, not if, these features become available to the corporate enterprise.

References

More Stories By Tom Petrocelli

Tom Petrocelli, president of Technology Alignment Partners, is a veteran of over 21 years in the technology arena. His background encompasses software engineering, marketing, IT, sales, marketing, and general management. He has worked in various industries including defense, digital signal processing, call center/CRM, networking, and data storage and storage networking. Tom is also the author of a new book entitled Data Protection and Information Lifecycle Management, published by Prentice Hall.

Comments (1)

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


@ThingsExpo Stories
The Internet of Things will greatly expand the opportunities for data collection and new business models driven off of that data. In her session at Internet of @ThingsExpo, Esmeralda Swartz, CMO of MetraTech, will discuss how for this to be effective you not only need to have infrastructure and operational models capable of utilizing this new phenomenon, but increasingly service providers will need to convince a skeptical public to participate. Get ready to show them the money! Speaker Bio: Esmeralda Swartz, CMO of MetraTech, has spent 16 years as a marketing, product management, and busin...
Samsung VP Jacopo Lenzi, who headed the company's recent SmartThings acquisition under the auspices of Samsung's Open Innovaction Center (OIC), answered a few questions we had about the deal. This interview was in conjunction with our interview with SmartThings CEO Alex Hawkinson. IoT Journal: SmartThings was developed in an open, standards-agnostic platform, and will now be part of Samsung's Open Innovation Center. Can you elaborate on your commitment to keep the platform open? Jacopo Lenzi: Samsung recognizes that true, accelerated innovation cannot be driven from one source, but requires a...
SYS-CON Events announced today that Red Hat, the world's leading provider of open source solutions, will exhibit at Internet of @ThingsExpo, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Red Hat is the world's leading provider of open source software solutions, using a community-powered approach to reliable and high-performing cloud, Linux, middleware, storage and virtualization technologies. Red Hat also offers award-winning support, training, and consulting services. As the connective hub in a global network of enterprises, partners, a...
P2P RTC will impact the landscape of communications, shifting from traditional telephony style communications models to OTT (Over-The-Top) cloud assisted & PaaS (Platform as a Service) communication services. The P2P shift will impact many areas of our lives, from mobile communication, human interactive web services, RTC and telephony infrastructure, user federation, security and privacy implications, business costs, and scalability. In his session at Internet of @ThingsExpo, Robin Raymond, Chief Architect at Hookflash Inc., will walk through the shifting landscape of traditional telephone a...
SYS-CON Events announced today that Matrix.org has been named “Silver Sponsor” of Internet of @ThingsExpo, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Matrix is an ambitious new open standard for open, distributed, real-time communication over IP. It defines a new approach for interoperable Instant Messaging and VoIP based on pragmatic HTTP APIs and WebRTC, and provides open source reference implementations to showcase and bootstrap the new standard. Our focus is on simplicity, security, and supporting the fullest feature set.
BSQUARE is a global leader of embedded software solutions. We enable smart connected systems at the device level and beyond that millions use every day and provide actionable data solutions for the growing Internet of Things (IoT) market. We empower our world-class customers with our products, services and solutions to achieve innovation and success. For more information, visit www.bsquare.com.
How do APIs and IoT relate? The answer is not as simple as merely adding an API on top of a dumb device, but rather about understanding the architectural patterns for implementing an IoT fabric. There are typically two or three trends: Exposing the device to a management framework Exposing that management framework to a business centric logic • Exposing that business layer and data to end users. This last trend is the IoT stack, which involves a new shift in the separation of what stuff happens, where data lives and where the interface lies. For instance, it’s a mix of architectural style...
From a software development perspective IoT is about programming "things," about connecting them with each other or integrating them with existing applications. In his session at @ThingsExpo, Yakov Fain, co-founder of Farata Systems and SuranceBay, will show you how small IoT-enabled devices from multiple manufacturers can be integrated into the workflow of an enterprise application. This is a practical demo of building a framework and components in HTML/Java/Mobile technologies to serve as a platform that can integrate new devices as they become available on the market.
SYS-CON Events announced today that SOA Software, an API management leader, will exhibit at SYS-CON's 15th International Cloud Expo®, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. SOA Software is a leading provider of API Management and SOA Governance products that equip business to deliver APIs and SOA together to drive their company to meet its business strategy quickly and effectively. SOA Software’s technology helps businesses to accelerate their digital channels with APIs, drive partner adoption, monetize their assets, and achieve a...
SYS-CON Events announced today that Utimaco will exhibit at SYS-CON's 15th International Cloud Expo®, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Utimaco is a leading manufacturer of hardware based security solutions that provide the root of trust to keep cryptographic keys safe, secure critical digital infrastructures and protect high value data assets. Only Utimaco delivers a general-purpose hardware security module (HSM) as a customizable platform to easily integrate into existing software solutions, embed business logic and build s...
Connected devices are changing the way we go about our everyday life, from wearables to driverless cars, to smart grids and entire industries revolutionizing business opportunities through smart objects, capable of two-way communication. But what happens when objects are given an IP-address, and we rely on that connection, sometimes with our lives? How do we secure those vast data infrastructures and safe-keep the privacy of sensitive information? This session will outline how each and every connected device can uphold a core root of trust via a unique cryptographic signature – a “bir...
Internet of @ThingsExpo Silicon Valley announced on Thursday its first 12 all-star speakers and sessions for its upcoming event, which will take place November 4-6, 2014, at the Santa Clara Convention Center in California. @ThingsExpo, the first and largest IoT event in the world, debuted at the Javits Center in New York City in June 10-12, 2014 with over 6,000 delegates attending the conference. Among the first 12 announced world class speakers, IBM will present two highly popular IoT sessions, which will take place November 4-6, 2014 at the Santa Clara Convention Center in Santa Clara, Calif...
Almost everyone sees the potential of Internet of Things but how can businesses truly unlock that potential. The key will be in the ability to discover business insight in the midst of an ocean of Big Data generated from billions of embedded devices via Systems of Discover. Businesses will also need to ensure that they can sustain that insight by leveraging the cloud for global reach, scale and elasticity.
WebRTC defines no default signaling protocol, causing fragmentation between WebRTC silos. SIP and XMPP provide possibilities, but come with considerable complexity and are not designed for use in a web environment. In his session at Internet of @ThingsExpo, Matthew Hodgson, technical co-founder of the Matrix.org, will discuss how Matrix is a new non-profit Open Source Project that defines both a new HTTP-based standard for VoIP & IM signaling and provides reference implementations.

SUNNYVALE, Calif., Oct. 20, 2014 /PRNewswire/ -- Spansion Inc. (NYSE: CODE), a global leader in embedded systems, today added 96 new products to the Spansion® FM4 Family of flexible microcontrollers (MCUs). Based on the ARM® Cortex®-M4F core, the new MCUs boast a 200 MHz operating frequency and support a diverse set of on-chip peripherals for enhanced human machine interfaces (HMIs) and machine-to-machine (M2M) communications. The rich set of periphera...

SYS-CON Events announced today that Aria Systems, the recurring revenue expert, has been named "Bronze Sponsor" of SYS-CON's 15th International Cloud Expo®, which will take place on November 4-6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Aria Systems helps leading businesses connect their customers with the products and services they love. Industry leaders like Pitney Bowes, Experian, AAA NCNU, VMware, HootSuite and many others choose Aria to power their recurring revenue business and deliver exceptional experiences to their customers.
The Internet of Things (IoT) is going to require a new way of thinking and of developing software for speed, security and innovation. This requires IT leaders to balance business as usual while anticipating for the next market and technology trends. Cloud provides the right IT asset portfolio to help today’s IT leaders manage the old and prepare for the new. Today the cloud conversation is evolving from private and public to hybrid. This session will provide use cases and insights to reinforce the value of the network in helping organizations to maximize their company’s cloud experience.
The Internet of Things (IoT) is making everything it touches smarter – smart devices, smart cars and smart cities. And lucky us, we’re just beginning to reap the benefits as we work toward a networked society. However, this technology-driven innovation is impacting more than just individuals. The IoT has an environmental impact as well, which brings us to the theme of this month’s #IoTuesday Twitter chat. The ability to remove inefficiencies through connected objects is driving change throughout every sector, including waste management. BigBelly Solar, located just outside of Boston, is trans...
SYS-CON Events announced today that Matrix.org has been named “Silver Sponsor” of Internet of @ThingsExpo, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Matrix is an ambitious new open standard for open, distributed, real-time communication over IP. It defines a new approach for interoperable Instant Messaging and VoIP based on pragmatic HTTP APIs and WebRTC, and provides open source reference implementations to showcase and bootstrap the new standard. Our focus is on simplicity, security, and supporting the fullest feature set.
Predicted by Gartner to add $1.9 trillion to the global economy by 2020, the Internet of Everything (IoE) is based on the idea that devices, systems and services will connect in simple, transparent ways, enabling seamless interactions among devices across brands and sectors. As this vision unfolds, it is clear that no single company can accomplish the level of interoperability required to support the horizontal aspects of the IoE. The AllSeen Alliance, announced in December 2013, was formed with the goal to advance IoE adoption and innovation in the connected home, healthcare, education, aut...