Click here to close now.




















Welcome!

Open Source Cloud Authors: David H Deans, Elizabeth White, Pat Romanski, Yakov Fain, Cloud Best Practices Network

Related Topics: Linux Containers, Open Source Cloud

Linux Containers: Article

Linux Cover Story — Multidimensional Tagging

Finding information naturally

Multidimensional tagging, a key component in social sharing sites, can potentially help enterprises manage large stores of information. In this article, I'll examine the ways that multidimensional tagging will be implemented using Open Source tools.

As storage costs have continued to decrease, organizations create and retain more information. It is easy - and useful - to keep information online, just in case it's needed. The challenge has become managing information not simply storing it. Structured systems, such as relational database management systems (RDBMS), have well-developed tools for indexing, locating, and retrieving information. Unfortunately not all information fits well into structured systems. Graphics, presentations, spreadsheets, and word processor documents have all proliferated, no longer bound by storage costs. This is especially true of desktop systems where 200GB hard drives are common.

Finding exactly the information needed at the right moment has become difficult. A typical desktop drive has tens of thousands of files. A file server or NAS device could easily have hundreds of thousands or even millions of files. Even if system files and applications are discounted, that's still a lot of files to wade through.

File Systems and Search Engines Provide Some Relief
File systems impose a certain degree of order on unstructured information. By allowing information to be placed in hierarchical directories, file systems group files that share some relationship into meaningful categories. The downside to categorization using a file system is complexity for the user. The user is forced to descend through more and more layers of directories to find what he wants. Just remembering where the correct directory is becomes a chore. Information organization via the file system becomes less useful when information is spread across an enterprise. Incompatible systems, individual ways of building directory structures, and shear scale make this a difficult way to manage information assets across an organization, even a small one.

Search engines provide some respite from our information organizational woes. By indexing keywords and storing references to their source in a database, it's possible to look quickly through all the available files. The vendors of searching engine technology, drawing from vast experience in indexing hundreds of millions of Web sites, provide tools that let users find information spread across a huge number of desktop disks and enterprise storage systems.

Tagging Provides Necessary Clues
The trouble with the typical search engine is that the user must first know what he's looking for. He has to have some idea what keywords are indexed for a particular piece of information. If the indexed words don't match the words that a user thinks apply to the information, then the search engine won't find what the user is looking for. A marketing brochure, for example, may not have any words in it that say "marketing brochure" but that's how the user thinks of it. The way the information is categorized in the users head doesn't always match the strict keyword index of the engine.

This has already become something of a problem for Web sites that allow users to share large amounts of information. Cues are needed to help visitors who come to a site find what interests them without knowing the exact nature of the content. In typical Internet fashion an organic solution has arisen called multidimensional tagging, labeling, social tagging, folksonomy, or just simply tagging. It's multidimensional because a single piece of information can have many different tags, reflecting the different dimensions that users apply to it. Tagging lets users assign a set of categories to a piece of information when they create it. The tag system, also known as the tag cloud, grows as people use the information and see relationships in it that the original author might have missed. Users categorize information according to how they view the information, which makes it useful for groups of people who don't always think alike, such as engineers and marketing people.

Tagging is a key feature in social sharing sites such as Yahoo's del.icio.us and Flickr as well as You Tube, and Userscripts.org. Whether it's sharing interesting Web pages, photos, video, or Greasemonkey scripts, all of these sites rely on user categorization. Without tagging, no one would find anything of interest and the site would fail. Unlike simple storage sites (such as Yahoo Photos and Yahoo Briefcase), they require a way of presenting information to users that lets them find it quickly. Users can find information even when they're not really sure what they want. Social sites don't abandon search engines. Instead they integrate searching with tagging to provide a breadth of information retrieval options. Most let users search the tag cloud as well as scanned keywords, providing a rich search environment.

Tagging: New to the Enterprise
Tagging technology for the enterprise environment is new and not widely deployed in products. That's unfortunate. Not only is it extremely useful for finding information, it's also a natural way to do it. This is especially true for people used to social sharing sites. The tagging methodology facilitates the efficient sharing of information across many users and a large file space. It's exactly what enterprises need to make best use of the great stores of unstructured information on their corporate networks.

There are three ways that tagging is being implemented in corporate environments: integrated into applications; as a part of a standalone information management system; and, eventually, as a file system feature. The first kind of implementation is readily available. Image management systems, even ones directed at the desktop environment like Google's Picasa, include tagging as a core feature. The next version of the Thunderbird e-mail client (version 2.0) is expected to include e-mail tagging, augmenting its current search capabilities. Of course, once users get used to tagging for managing certain types of information, they will wonder why they can't use it for all the information that they need to access. They'll expect tag clouds that span all kinds of information in the enterprise.

Tagging is also being implemented in targeted information management tools. Tools for searching large stores of information in a corporate network are still at an early stage but tagging should be expected to make an appearance in information management and search engine tools in the near future. Consider this, Yahoo uses tagging in Flickr and del.icio.us as well as the upcoming My Web 2.0. Is it a stretch to expect it to implement tagging in the corporate search arena? The same is true for Google, which uses tagging on its eBlogger site, GMail Web-based mail service, and Picasa image management tool.

Finally, tagging can be expected to become a feature of the file system and operating system. Some aspects of tagging already exist in operating systems. The ability to attach keywords to files in Microsoft Windows is an example of file system-level tagging. These keywords are currently read by Microsoft's desktop search engine, creating a crude multidimensional tagging feature. Of course, entering and displaying tags is clunky and tags can't be displayed in and of themselves, rendering it more of a hack than a real feature. It does, however, point the way to future features of the operating system. Fully integrated into an operating system as normal metadata and using standard visual cues such those used on social sites, tagging will become a typical part of most corporate environments.

Tools Exist, File System Hooks Don't
The tools for corporate tagging capabilities already exist in the Open Source community. Most of it is encapsulated in the tools used by social bookmarking sites, which are often based on the LAMP stack. They're typically written in common scripting languages, such as Perl or Python, or Java. One such Open Source tool is unalog. Ostensibly a social bookmarking system, it's written in Python and the source is readily available on SourceForge. While the core tools exist, the hooks into the file system are still mostly missing.

A somewhat different but innovative approach is evident with Flickrfs or the Flickr File System. Based on FUSE, it creates a virtual file system with tagging for the Flickr digital photo management service. A fusion of file system and service, Flickrfs lets Linux users access the Flickr service as if it were any other mounted Linux file system. Photos can be accessed through the same tags available on Flickr using standard Linux commands such as cp. Flickrfs represents another way that tagging may come to information management - as a specific application or service but integrated into the normal file system.

Conclusion
Multidimensional tagging provides an opportunity to let users manage information more in line with their natural way of thinking. By sharing tags across the enterprise, users will spend less time looking for information and more time making use of it. Unlike other collaborative systems, users do all the work without legions of editors making decisions that users find mystifying. The social sites on the Internet have shown this to be a viable information management model. It's a matter of how and when, not if, these features become available to the corporate enterprise.

References

More Stories By Tom Petrocelli

Tom Petrocelli, president of Technology Alignment Partners, is a veteran of over 21 years in the technology arena. His background encompasses software engineering, marketing, IT, sales, marketing, and general management. He has worked in various industries including defense, digital signal processing, call center/CRM, networking, and data storage and storage networking. Tom is also the author of a new book entitled Data Protection and Information Lifecycle Management, published by Prentice Hall.

Comments (1) View Comments

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


Most Recent Comments
Linux News Desk 05/21/06 09:33:34 AM EDT

Multidimensional tagging, a key component in social sharing sites, can potentially help enterprises manage large stores of information. In this article, I'll examine the ways that multidimensional tagging will be implemented using Open Source tools.

@ThingsExpo Stories
With the Apple Watch making its way onto wrists all over the world, it’s only a matter of time before it becomes a staple in the workplace. In fact, Forrester reported that 68 percent of technology and business decision-makers characterize wearables as a top priority for 2015. Recognizing their business value early on, FinancialForce.com was the first to bring ERP to wearables, helping streamline communication across front and back office functions. In his session at @ThingsExpo, Kevin Roberts, GM of Platform at FinancialForce.com, will discuss the value of business applications on wearable ...
Contrary to mainstream media attention, the multiple possibilities of how consumer IoT will transform our everyday lives aren’t the only angle of this headline-gaining trend. There’s a huge opportunity for “industrial IoT” and “Smart Cities” to impact the world in the same capacity – especially during critical situations. For example, a community water dam that needs to release water can leverage embedded critical communications logic to alert the appropriate individuals, on the right device, as soon as they are needed to take action.
SYS-CON Events announced today that HPM Networks will exhibit at the 17th International Cloud Expo®, which will take place on November 3–5, 2015, at the Santa Clara Convention Center in Santa Clara, CA. For 20 years, HPM Networks has been integrating technology solutions that solve complex business challenges. HPM Networks has designed solutions for both SMB and enterprise customers throughout the San Francisco Bay Area.
SYS-CON Events announced today that Micron Technology, Inc., a global leader in advanced semiconductor systems, will exhibit at the 17th International Cloud Expo®, which will take place on November 3–5, 2015, at the Santa Clara Convention Center in Santa Clara, CA. Micron’s broad portfolio of high-performance memory technologies – including DRAM, NAND and NOR Flash – is the basis for solid state drives, modules, multichip packages and other system solutions. Backed by more than 35 years of technology leadership, Micron's memory solutions enable the world's most innovative computing, consumer,...
As more intelligent IoT applications shift into gear, they’re merging into the ever-increasing traffic flow of the Internet. It won’t be long before we experience bottlenecks, as IoT traffic peaks during rush hours. Organizations that are unprepared will find themselves by the side of the road unable to cross back into the fast lane. As billions of new devices begin to communicate and exchange data – will your infrastructure be scalable enough to handle this new interconnected world?
Through WebRTC, audio and video communications are being embedded more easily than ever into applications, helping carriers, enterprises and independent software vendors deliver greater functionality to their end users. With today’s business world increasingly focused on outcomes, users’ growing calls for ease of use, and businesses craving smarter, tighter integration, what’s the next step in delivering a richer, more immersive experience? That richer, more fully integrated experience comes about through a Communications Platform as a Service which allows for messaging, screen sharing, video...
SYS-CON Events announced today that Pythian, a global IT services company specializing in helping companies leverage disruptive technologies to optimize revenue-generating systems, has been named “Bronze Sponsor” of SYS-CON's 17th Cloud Expo, which will take place on November 3–5, 2015, at the Santa Clara Convention Center in Santa Clara, CA. Founded in 1997, Pythian is a global IT services company that helps companies compete by adopting disruptive technologies such as cloud, Big Data, advanced analytics, and DevOps to advance innovation and increase agility. Specializing in designing, imple...
In his session at @ThingsExpo, Lee Williams, a producer of the first smartphones and tablets, will talk about how he is now applying his experience in mobile technology to the design and development of the next generation of Environmental and Sustainability Services at ETwater. He will explain how M2M controllers work through wirelessly connected remote controls; and specifically delve into a retrofit option that reverse-engineers control codes of existing conventional controller systems so they don't have to be replaced and are instantly converted to become smart, connected devices.
SYS-CON Events announced today that IceWarp will exhibit at the 17th International Cloud Expo®, which will take place on November 3–5, 2015, at the Santa Clara Convention Center in Santa Clara, CA. IceWarp, the leader of cloud and on-premise messaging, delivers secured email, chat, documents, conferencing and collaboration to today's mobile workforce, all in one unified interface
WebRTC has had a real tough three or four years, and so have those working with it. Only a few short years ago, the development world were excited about WebRTC and proclaiming how awesome it was. You might have played with the technology a couple of years ago, only to find the extra infrastructure requirements were painful to implement and poorly documented. This probably left a bitter taste in your mouth, especially when things went wrong.
Too often with compelling new technologies market participants become overly enamored with that attractiveness of the technology and neglect underlying business drivers. This tendency, what some call the “newest shiny object syndrome,” is understandable given that virtually all of us are heavily engaged in technology. But it is also mistaken. Without concrete business cases driving its deployment, IoT, like many other technologies before it, will fade into obscurity.
While many app developers are comfortable building apps for the smartphone, there is a whole new world out there. In his session at @ThingsExpo, Narayan Sainaney, Co-founder and CTO of Mojio, will discuss how the business case for connected car apps is growing and, with open platform companies having already done the heavy lifting, there really is no barrier to entry.
Consumer IoT applications provide data about the user that just doesn’t exist in traditional PC or mobile web applications. This rich data, or “context,” enables the highly personalized consumer experiences that characterize many consumer IoT apps. This same data is also providing brands with unprecedented insight into how their connected products are being used, while, at the same time, powering highly targeted engagement and marketing opportunities. In his session at @ThingsExpo, Nathan Treloar, President and COO of Bebaio, will explore examples of brands transforming their businesses by t...
With the proliferation of connected devices underpinning new Internet of Things systems, Brandon Schulz, Director of Luxoft IoT – Retail, will be looking at the transformation of the retail customer experience in brick and mortar stores in his session at @ThingsExpo. Questions he will address include: Will beacons drop to the wayside like QR codes, or be a proximity-based profit driver? How will the customer experience change in stores of all types when everything can be instrumented and analyzed? As an area of investment, how might a retail company move towards an innovation methodolo...
The Internet of Things (IoT) is about the digitization of physical assets including sensors, devices, machines, gateways, and the network. It creates possibilities for significant value creation and new revenue generating business models via data democratization and ubiquitous analytics across IoT networks. The explosion of data in all forms in IoT requires a more robust and broader lens in order to enable smarter timely actions and better outcomes. Business operations become the key driver of IoT applications and projects. Business operations, IT, and data scientists need advanced analytics t...
As more and more data is generated from a variety of connected devices, the need to get insights from this data and predict future behavior and trends is increasingly essential for businesses. Real-time stream processing is needed in a variety of different industries such as Manufacturing, Oil and Gas, Automobile, Finance, Online Retail, Smart Grids, and Healthcare. Azure Stream Analytics is a fully managed distributed stream computation service that provides low latency, scalable processing of streaming data in the cloud with an enterprise grade SLA. It features built-in integration with Azur...
Akana has announced the availability of the new Akana Healthcare Solution. The API-driven solution helps healthcare organizations accelerate their transition to being secure, digitally interoperable businesses. It leverages the Health Level Seven International Fast Healthcare Interoperability Resources (HL7 FHIR) standard to enable broader business use of medical data. Akana developed the Healthcare Solution in response to healthcare businesses that want to increase electronic, multi-device access to health records while reducing operating costs and complying with government regulations.
For IoT to grow as quickly as analyst firms’ project, a lot is going to fall on developers to quickly bring applications to market. But the lack of a standard development platform threatens to slow growth and make application development more time consuming and costly, much like we’ve seen in the mobile space. In his session at @ThingsExpo, Mike Weiner, Product Manager of the Omega DevCloud with KORE Telematics Inc., discussed the evolving requirements for developers as IoT matures and conducted a live demonstration of how quickly application development can happen when the need to comply wit...
The Internet of Everything (IoE) brings together people, process, data and things to make networked connections more relevant and valuable than ever before – transforming information into knowledge and knowledge into wisdom. IoE creates new capabilities, richer experiences, and unprecedented opportunities to improve business and government operations, decision making and mission support capabilities.
Explosive growth in connected devices. Enormous amounts of data for collection and analysis. Critical use of data for split-second decision making and actionable information. All three are factors in making the Internet of Things a reality. Yet, any one factor would have an IT organization pondering its infrastructure strategy. How should your organization enhance its IT framework to enable an Internet of Things implementation? In his session at @ThingsExpo, James Kirkland, Red Hat's Chief Architect for the Internet of Things and Intelligent Systems, described how to revolutionize your archit...