Internal Information Search Using Search Engines—Defining the Problem

Rapidly growing volumes of data, plus countless communication channels such as chat, email, wikis, and documents. Can enterprise search engines help here?

Reading duration: approx. 3 Minutes

We have long been in an era of unchecked data growth. The nearly limitless tracking of every interaction, every activity, and every process is generating more and more data.
The IDC (International Data Corporation) estimates that by 2025, a total of 175 zettabytes of data will be generated worldwide.

| 1 zettabyte (1ZB) = 10^21 bytes
| If stacked on top of each other, 175 zettabytes would reach 2.31 times as far as the Moon

The data is not limited to IoT telemetry relevant to big data or measurement data from production lines: Additional parallel sources of information are generated by the growing number of internal communication channels, as well as by efforts to consolidate and aggregate resources. The lack of streamlining processes for internal resources, structures, and knowledge management tools leads to a steady increase in the volume of data within companies.

Employees spend hours searching for information

Employees spend a substantial portion of their workweek—about 8.8 hours—searching for information in this data, whether it’s the port for a specific service or their boss’s internal extension [9]. Not all systems support a comprehensive “Google-style” search; some don’t offer any search functionality at all. Where search engines do exist, they vary greatly in their functionality and scope—tools like Jira, for example, offer a query language, while some chat platforms don’t even allow cross-channel searches. If no search engine is available, employees must wade through folder structures and page hierarchies and have at least a rough idea of where the information they’re looking for is located.

The Promise of Enterprise Search

The field of enterprise search addresses this issue and aims to provide a single, comprehensive search solution within the company that can meet the information needs of all employees. Commercial providers such as Splunk and Amazon offer cloud-based solutions, but open-source software like Elasticsearch and Solr can also be used for enterprise search.
The promise is simple: using Enterprise Search is expected to significantly reduce the time employees spend searching for data. To achieve this, all data sources are made searchable with a single query.

The providers promise high returns on investment:
Google uses “conservative figures” [4, p. 4] of over 22 million euros for medium-sized companies [4], while John Lenker of LucidWorks estimates an annual ROI of 15 to 20 times the investment [5].

It's more complex than on the web

However, it is questionable whether these figures are accurate. These are generally rough estimates, and in most cases, potentially high training and support costs are not taken into account at all.
The results of the few scientific studies on the cost-effectiveness and effectiveness of enterprise search are sobering: Company-wide search capabilities quickly reach their limits due to repetitive patterns: The larger a company is and the more diverse its silos of knowledge, the less efficient a search across all information can become. Other soft factors include a company’s structures and hierarchies, the presence of a basic understanding of search, and the quality and preparation of the underlying data [6].

Searching an intranet or internal databases is significantly more complex than searching the web:
Relevant data comes from various sources in a wide variety of formats; it may be fragmented or incomplete and is only aggregated during the review process. In addition, individual and fine-grained authorization models must be taken into account. Data protection requirements further increase the effort involved.
The searcher expects not only that various structured data be made searchable, but also that they find the correct result—not merely the best possible one [7,8].

It also depends on the employees

Further problems arise from employees’ search habits and behavior.
Intranet searchers can be divided into three groups, whose needs are in some cases incompatible (e.g., recall vs. precision [10]) [11]: inexperienced users and occasional searchers (approx. 80%), interactive users and heavy users (approx. 14%), and “employees skilled in information retrieval,” . To reduce the cognitive load, users have developed various strategies. For example, some navigate through folder and page hierarchies to avoid formulating a complex query that would have allowed them to jump directly to the result.

The current hype surrounding enterprise search should be viewed with caution. When calculating ROI, one must consider not only potential costs but also the suitability of the company’s internal structures and the competencies of its employees.

Sources:

[1] https://www.seagate.com/files/www-content/our-story/trends/files/idc-seagate-dataage-whitepaper.pdf
[2] https://sci-hub.se/10.1145/2637748.2638425
[3] https://sci-hub.se/10.1145/985692.985745
[4] https://static.googleusercontent.com/media/194.78.99.204/en/204/enterprise/search/files/Internal_Search_ROI.pdf
[5] https://de.slideshare.net/lucidworks/measuring-roi-on-enterprise-search-john-lenker-lucidworks
[6] https://edoc.hu-berlin.de/bitstream/handle/18452/17001/bertram.pdf?sequence=1&isAllowed=y
[7] Mani Abrol, Neil Latarche, Uma Mahadevan, Jianchang Mao, Rajat Mukherjee, Prabhakar Raghavan, Michel Tourn, John Wang, and Grace Zhang. Navigating large-scale semi-structured data in business portals. Very Large Databases, pp. 663–666, 2001.
[8] Rajat Mukherjee and Jianchang Mao. Enterprise Search: Tough Stuff. ACM Queue, 2(2):36–46, 2004.
[9] McKinsey Global Institute. The Social Economy: Unlocking Value and Productivity through Social Technologies. 2012.
[10] https://de.wikipedia.org/wiki/Beurteilung_eines_bin%C3%A4ren_Klassifikators#Anwendung_im_Information_Retrieva
[11] Dick Stenmark. Identifying clusters of user behavior in intranet search engine log files. Journal of the American Society for Information Science and Technology, 59(14):2232–2243, Dec. 2008. doi:10.1002/asi.20931.

Share:

More articles

Wer will, findet Wege, wer nicht will, findet Gründe
Frank Keller, Product Owner / Digital Consultant at punkt.de
Working at punkt.de