Welcome!

Open Source Cloud Authors: Pat Romanski, Liz McMillan, Stackify Blog, Wesley Coelho, Rainer Ersch

Related Topics: Linux Containers, Open Source Cloud

Linux Containers: Article

Linux Cover Story — Multidimensional Tagging

Finding information naturally

Multidimensional tagging, a key component in social sharing sites, can potentially help enterprises manage large stores of information. In this article, I'll examine the ways that multidimensional tagging will be implemented using Open Source tools.

As storage costs have continued to decrease, organizations create and retain more information. It is easy - and useful - to keep information online, just in case it's needed. The challenge has become managing information not simply storing it. Structured systems, such as relational database management systems (RDBMS), have well-developed tools for indexing, locating, and retrieving information. Unfortunately not all information fits well into structured systems. Graphics, presentations, spreadsheets, and word processor documents have all proliferated, no longer bound by storage costs. This is especially true of desktop systems where 200GB hard drives are common.

Finding exactly the information needed at the right moment has become difficult. A typical desktop drive has tens of thousands of files. A file server or NAS device could easily have hundreds of thousands or even millions of files. Even if system files and applications are discounted, that's still a lot of files to wade through.

File Systems and Search Engines Provide Some Relief
File systems impose a certain degree of order on unstructured information. By allowing information to be placed in hierarchical directories, file systems group files that share some relationship into meaningful categories. The downside to categorization using a file system is complexity for the user. The user is forced to descend through more and more layers of directories to find what he wants. Just remembering where the correct directory is becomes a chore. Information organization via the file system becomes less useful when information is spread across an enterprise. Incompatible systems, individual ways of building directory structures, and shear scale make this a difficult way to manage information assets across an organization, even a small one.

Search engines provide some respite from our information organizational woes. By indexing keywords and storing references to their source in a database, it's possible to look quickly through all the available files. The vendors of searching engine technology, drawing from vast experience in indexing hundreds of millions of Web sites, provide tools that let users find information spread across a huge number of desktop disks and enterprise storage systems.

Tagging Provides Necessary Clues
The trouble with the typical search engine is that the user must first know what he's looking for. He has to have some idea what keywords are indexed for a particular piece of information. If the indexed words don't match the words that a user thinks apply to the information, then the search engine won't find what the user is looking for. A marketing brochure, for example, may not have any words in it that say "marketing brochure" but that's how the user thinks of it. The way the information is categorized in the users head doesn't always match the strict keyword index of the engine.

This has already become something of a problem for Web sites that allow users to share large amounts of information. Cues are needed to help visitors who come to a site find what interests them without knowing the exact nature of the content. In typical Internet fashion an organic solution has arisen called multidimensional tagging, labeling, social tagging, folksonomy, or just simply tagging. It's multidimensional because a single piece of information can have many different tags, reflecting the different dimensions that users apply to it. Tagging lets users assign a set of categories to a piece of information when they create it. The tag system, also known as the tag cloud, grows as people use the information and see relationships in it that the original author might have missed. Users categorize information according to how they view the information, which makes it useful for groups of people who don't always think alike, such as engineers and marketing people.

Tagging is a key feature in social sharing sites such as Yahoo's del.icio.us and Flickr as well as You Tube, and Userscripts.org. Whether it's sharing interesting Web pages, photos, video, or Greasemonkey scripts, all of these sites rely on user categorization. Without tagging, no one would find anything of interest and the site would fail. Unlike simple storage sites (such as Yahoo Photos and Yahoo Briefcase), they require a way of presenting information to users that lets them find it quickly. Users can find information even when they're not really sure what they want. Social sites don't abandon search engines. Instead they integrate searching with tagging to provide a breadth of information retrieval options. Most let users search the tag cloud as well as scanned keywords, providing a rich search environment.

Tagging: New to the Enterprise
Tagging technology for the enterprise environment is new and not widely deployed in products. That's unfortunate. Not only is it extremely useful for finding information, it's also a natural way to do it. This is especially true for people used to social sharing sites. The tagging methodology facilitates the efficient sharing of information across many users and a large file space. It's exactly what enterprises need to make best use of the great stores of unstructured information on their corporate networks.

There are three ways that tagging is being implemented in corporate environments: integrated into applications; as a part of a standalone information management system; and, eventually, as a file system feature. The first kind of implementation is readily available. Image management systems, even ones directed at the desktop environment like Google's Picasa, include tagging as a core feature. The next version of the Thunderbird e-mail client (version 2.0) is expected to include e-mail tagging, augmenting its current search capabilities. Of course, once users get used to tagging for managing certain types of information, they will wonder why they can't use it for all the information that they need to access. They'll expect tag clouds that span all kinds of information in the enterprise.

Tagging is also being implemented in targeted information management tools. Tools for searching large stores of information in a corporate network are still at an early stage but tagging should be expected to make an appearance in information management and search engine tools in the near future. Consider this, Yahoo uses tagging in Flickr and del.icio.us as well as the upcoming My Web 2.0. Is it a stretch to expect it to implement tagging in the corporate search arena? The same is true for Google, which uses tagging on its eBlogger site, GMail Web-based mail service, and Picasa image management tool.

Finally, tagging can be expected to become a feature of the file system and operating system. Some aspects of tagging already exist in operating systems. The ability to attach keywords to files in Microsoft Windows is an example of file system-level tagging. These keywords are currently read by Microsoft's desktop search engine, creating a crude multidimensional tagging feature. Of course, entering and displaying tags is clunky and tags can't be displayed in and of themselves, rendering it more of a hack than a real feature. It does, however, point the way to future features of the operating system. Fully integrated into an operating system as normal metadata and using standard visual cues such those used on social sites, tagging will become a typical part of most corporate environments.

Tools Exist, File System Hooks Don't
The tools for corporate tagging capabilities already exist in the Open Source community. Most of it is encapsulated in the tools used by social bookmarking sites, which are often based on the LAMP stack. They're typically written in common scripting languages, such as Perl or Python, or Java. One such Open Source tool is unalog. Ostensibly a social bookmarking system, it's written in Python and the source is readily available on SourceForge. While the core tools exist, the hooks into the file system are still mostly missing.

A somewhat different but innovative approach is evident with Flickrfs or the Flickr File System. Based on FUSE, it creates a virtual file system with tagging for the Flickr digital photo management service. A fusion of file system and service, Flickrfs lets Linux users access the Flickr service as if it were any other mounted Linux file system. Photos can be accessed through the same tags available on Flickr using standard Linux commands such as cp. Flickrfs represents another way that tagging may come to information management - as a specific application or service but integrated into the normal file system.

Conclusion
Multidimensional tagging provides an opportunity to let users manage information more in line with their natural way of thinking. By sharing tags across the enterprise, users will spend less time looking for information and more time making use of it. Unlike other collaborative systems, users do all the work without legions of editors making decisions that users find mystifying. The social sites on the Internet have shown this to be a viable information management model. It's a matter of how and when, not if, these features become available to the corporate enterprise.

References

More Stories By Tom Petrocelli

Tom Petrocelli, president of Technology Alignment Partners, is a veteran of over 21 years in the technology arena. His background encompasses software engineering, marketing, IT, sales, marketing, and general management. He has worked in various industries including defense, digital signal processing, call center/CRM, networking, and data storage and storage networking. Tom is also the author of a new book entitled Data Protection and Information Lifecycle Management, published by Prentice Hall.

Comments (1) View Comments

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


Most Recent Comments
Linux News Desk 05/21/06 09:33:34 AM EDT

Multidimensional tagging, a key component in social sharing sites, can potentially help enterprises manage large stores of information. In this article, I'll examine the ways that multidimensional tagging will be implemented using Open Source tools.

@ThingsExpo Stories
Organizations do not need a Big Data strategy; they need a business strategy that incorporates Big Data. Most organizations lack a road map for using Big Data to optimize key business processes, deliver a differentiated customer experience, or uncover new business opportunities. They do not understand what’s possible with respect to integrating Big Data into the business model.
Enterprises have taken advantage of IoT to achieve important revenue and cost advantages. What is less apparent is how incumbent enterprises operating at scale have, following success with IoT, built analytic, operations management and software development capabilities – ranging from autonomous vehicles to manageable robotics installations. They have embraced these capabilities as if they were Silicon Valley startups. As a result, many firms employ new business models that place enormous impor...
Amazon is pursuing new markets and disrupting industries at an incredible pace. Almost every industry seems to be in its crosshairs. Companies and industries that once thought they were safe are now worried about being “Amazoned.”. The new watch word should be “Be afraid. Be very afraid.” In his session 21st Cloud Expo, Chris Kocher, a co-founder of Grey Heron, will address questions such as: What new areas is Amazon disrupting? How are they doing this? Where are they likely to go? What are th...
SYS-CON Events announced today that MIRAI Inc. will exhibit at the Japan External Trade Organization (JETRO) Pavilion at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. MIRAI Inc. are IT consultants from the public sector whose mission is to solve social issues by technology and innovation and to create a meaningful future for people.
SYS-CON Events announced today that Dasher Technologies will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Dasher Technologies, Inc. ® is a premier IT solution provider that delivers expert technical resources along with trusted account executives to architect and deliver complete IT solutions and services to help our clients execute their goals, plans and objectives. Since 1999, we'v...
SYS-CON Events announced today that NetApp has been named “Bronze Sponsor” of SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. NetApp is the data authority for hybrid cloud. NetApp provides a full range of hybrid cloud data services that simplify management of applications and data across cloud and on-premises environments to accelerate digital transformation. Together with their partners, NetApp emp...
SYS-CON Events announced today that IBM has been named “Diamond Sponsor” of SYS-CON's 21st Cloud Expo, which will take place on October 31 through November 2nd 2017 at the Santa Clara Convention Center in Santa Clara, California.
SYS-CON Events announced today that TidalScale, a leading provider of systems and services, will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. TidalScale has been involved in shaping the computing landscape. They've designed, developed and deployed some of the most important and successful systems and services in the history of the computing industry - internet, Ethernet, operating s...
Infoblox delivers Actionable Network Intelligence to enterprise, government, and service provider customers around the world. They are the industry leader in DNS, DHCP, and IP address management, the category known as DDI. We empower thousands of organizations to control and secure their networks from the core-enabling them to increase efficiency and visibility, improve customer service, and meet compliance requirements.
SYS-CON Events announced today that IBM has been named “Diamond Sponsor” of SYS-CON's 21st Cloud Expo, which will take place on October 31 through November 2nd 2017 at the Santa Clara Convention Center in Santa Clara, California.
Join IBM November 1 at 21st Cloud Expo at the Santa Clara Convention Center in Santa Clara, CA, and learn how IBM Watson can bring cognitive services and AI to intelligent, unmanned systems. Cognitive analysis impacts today’s systems with unparalleled ability that were previously available only to manned, back-end operations. Thanks to cloud processing, IBM Watson can bring cognitive services and AI to intelligent, unmanned systems. Imagine a robot vacuum that becomes your personal assistant tha...
In a recent survey, Sumo Logic surveyed 1,500 customers who employ cloud services such as Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP). According to the survey, a quarter of the respondents have already deployed Docker containers and nearly as many (23 percent) are employing the AWS Lambda serverless computing framework. It’s clear: serverless is here to stay. The adoption does come with some needed changes, within both application development and operations. Tha...
In his Opening Keynote at 21st Cloud Expo, John Considine, General Manager of IBM Cloud Infrastructure, will lead you through the exciting evolution of the cloud. He'll look at this major disruption from the perspective of technology, business models, and what this means for enterprises of all sizes. John Considine is General Manager of Cloud Infrastructure Services at IBM. In that role he is responsible for leading IBM’s public cloud infrastructure including strategy, development, and offering ...
SYS-CON Events announced today that Avere Systems, a leading provider of enterprise storage for the hybrid cloud, will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 - Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Avere delivers a more modern architectural approach to storage that doesn't require the overprovisioning of storage capacity to achieve performance, overspending on expensive storage media for inactive data or the overbui...
Widespread fragmentation is stalling the growth of the IIoT and making it difficult for partners to work together. The number of software platforms, apps, hardware and connectivity standards is creating paralysis among businesses that are afraid of being locked into a solution. EdgeX Foundry is unifying the community around a common IoT edge framework and an ecosystem of interoperable components.
SYS-CON Events announced today that TidalScale will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. TidalScale is the leading provider of Software-Defined Servers that bring flexibility to modern data centers by right-sizing servers on the fly to fit any data set or workload. TidalScale’s award-winning inverse hypervisor technology combines multiple commodity servers (including their ass...
SYS-CON Events announced today that N3N will exhibit at SYS-CON's @ThingsExpo, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. N3N’s solutions increase the effectiveness of operations and control centers, increase the value of IoT investments, and facilitate real-time operational decision making. N3N enables operations teams with a four dimensional digital “big board” that consolidates real-time live video feeds alongside IoT sensor data a...
As hybrid cloud becomes the de-facto standard mode of operation for most enterprises, new challenges arise on how to efficiently and economically share data across environments. In his session at 21st Cloud Expo, Dr. Allon Cohen, VP of Product at Elastifile, will explore new techniques and best practices that help enterprise IT benefit from the advantages of hybrid cloud environments by enabling data availability for both legacy enterprise and cloud-native mission critical applications. By rev...
With major technology companies and startups seriously embracing Cloud strategies, now is the perfect time to attend 21st Cloud Expo October 31 - November 2, 2017, at the Santa Clara Convention Center, CA, and June 12-14, 2018, at the Javits Center in New York City, NY, and learn what is going on, contribute to the discussions, and ensure that your enterprise is on the right path to Digital Transformation.
Join IBM November 1 at 21st Cloud Expo at the Santa Clara Convention Center in Santa Clara, CA, and learn how IBM Watson can bring cognitive services and AI to intelligent, unmanned systems. Cognitive analysis impacts today’s systems with unparalleled ability that were previously available only to manned, back-end operations. Thanks to cloud processing, IBM Watson can bring cognitive services and AI to intelligent, unmanned systems. Imagine a robot vacuum that becomes your personal assistant th...