具体描述
"The authors, the best minds on the topic, are breaking new ground. They show how every organization can realize the benefits of a system that can search and present complex ideas or data from what has been a mostly untapped source of raw data." --Randy Chalfant, CTO, Sun Microsystems The Definitive Guide to Unstructured Data Management and Analysis--From the World's Leading Information Management Expert A wealth of invaluable information exists in unstructured textual form, but organizations have found it difficult or impossible to access and utilize it. This is changing rapidly: new approaches finally make it possible to glean useful knowledge from virtually any collection of unstructured data. William H. Inmon--the father of data warehousing--and Anthony Nesavich introduce the next data revolution: unstructured data management. Inmon and Nesavich cover all you need to know to make unstructured data work for your organization. You'll learn how to bring it into your existing structured data environment, leverage existing analytical infrastructure, and implement textual analytic processing technologies to solve new problems and uncover new opportunities. Inmon and Nesavich introduce breakthrough techniques covered in no other book--including the powerful role of textual integration, new ways to integrate textual data into data warehouses, and new SQL techniques for reading and analyzing text. They also present five chapter-length, real-world case studies--demonstrating unstructured data at work in medical research, insurance, chemical manufacturing, contracting, and beyond. This book will be indispensable to every business and technical professional trying to make sense of a large body of unstructured text: managers, database designers, data modelers, DBAs, researchers, and end users alike. Coverage includes *What unstructured data is, and how it differs from structured data*First generation technology for handling unstructured data, from search engines to ECM--and its limitations*Integrating text so it can be analyzed with a common, colloquial vocabulary: integration engines, ontologies, glossaries, and taxonomies*Processing semistructured data: uncovering patterns, words, identifiers, and conflicts*Novel processing opportunities that arise when text is freed from context *Architecture and unstructured data: Data Warehousing 2.0 *Building unstructured relational databases and linking them to structured data*Visualizations and Self-Organizing Maps (SOMs), including Compudigm and Raptor solutions*Capturing knowledge from spreadsheet data and email*Implementing and managing metadata: data models, data quality, and more William H. Inmon is founder, president, and CTO of Inmon Data Systems. He is the father of the data warehouse concept, the corporate information factory, and the government information factory. Inmon has written 47 books on data warehouse, database, and information technology management; as well as more than 750 articles for trade journals such as Data Management Review, Byte, Datamation, and ComputerWorld. His b-eye-network.com newsletter currently reaches 55,000 people. Anthony Nesavich worked at Inmon Data Systems, where he developed multiple reports that successfully query unstructured data. Preface xvii 1 Unstructured Textual Data in the Organization 1 2 The Environments of Structured Data and Unstructured Data 15 3 First Generation Textual Analytics 33 4 Integrating Unstructured Text into the Structured Environment 47 5 Semistructured Data 73 6 Architecture and Textual Analytics 83 7 The Unstructured Database 95 8 Analyzing a Combination of Unstructured Data and Structured Data 113 9 Analyzing Text Through Visualization 127 10 Spreadsheets and Email 135 11 Metadata in Unstructured Data 147 12 A Methodology for Textual Analytics 163 13 Merging Unstructured Databases into the Data Warehouse 175 14 Using SQL to Analyze Text 185 15 Case Study--Textual Analytics in Medical Research 195 16 Case Study--A Database for Harmful Chemicals 203 17 Case Study--Managing Contracts Through an Unstructured Database 209 18 Case Study--Creating a Corporate Taxonomy (Glossary) 215 19 Case Study--Insurance Claims 219 Glossary 227 Index 233
作者简介
目录信息
读后感
用户评价
这本书的行文风格简直是一股清流,它充满了知识分子的批判精神,同时又保持着一种令人愉悦的亲和力。阅读体验非常流畅,仿佛在和一位睿智的长者进行一次深入的咖啡馆对话。作者在叙述中展现出的那种深厚的人文素养,使得原本可能枯燥的技术讨论,变得富有哲理和诗意。例如,他将信息的不完整性比喻为音乐中的休止符,强调了缺失数据本身可能携带的意义,这种比喻极具启发性。更令人赞叹的是,作者在全书中始终保持着一种开放的心态,他从不宣称自己找到了“唯一的真理”,而是不断鼓励读者去质疑、去实验、去构建属于自己的分析框架。我尤其欣赏作者在介绍新兴技术趋势时的审慎态度,他没有盲目追捧热点,而是用历史的眼光去审视这些新工具的长期潜力与潜在陷阱。这种冷静而深刻的分析,极大地提升了这本书的权威性和实用价值,它不是提供一个速成的解决方案,而是教授一种可持续的、面向未来的分析方法论。对于那些渴望超越当前技术栈,追求更高层次数据素养的专业人士来说,这本书无疑是一剂强心针。
我对这本书的阅读过程,是一次持续的“认知重塑”。我原本以为数据分析就是遵循既定的流程,但这本书教会我,流程本身就是一种需要不断被挑战的对象。作者在论述中构建了一个多维度的分析矩阵,帮助读者识别出不同类型数据源的固有结构缺陷,并据此设计出针对性的捕获和解读策略。特别是关于时间序列数据的非线性模式识别部分,作者提出的那套评估体系,极大地拓宽了我对复杂系统建模的思路。这本书的深度在于它敢于直面那些“难以量化”的部分——那些往往被传统量化工具所忽视的、充满人性弱点和系统误差的领域。书中大量的脚注和引文,清晰地勾勒出作者学术思想的来源和演变脉络,显示出其扎实的理论基础。阅读完这本书后,我感到自己的分析工具箱被彻底升级了,不再是简单地使用既有的螺丝刀和扳手,而是学会了如何冶炼和铸造最适合特定任务的工具。这种从“应用层”向“原理层”的提升,是任何速成读物都无法提供的宝贵财富。
这本书的装帧设计虽然内敛,但其内容的分量却远超预期。它不是那种读完后可以束之高阁的参考书,而是一本需要经常翻阅、时常在关键时刻提供灵感支撑的案头宝典。作者对于信息熵和复杂性的阐释,有着一种近乎物理学般的严谨和美感。他没有回避数据世界的混乱本质,反而将其视为创造力的源泉。我特别喜欢书中对“叙事驱动的分析”这一概念的深入探讨,这部分内容完美地衔接了技术分析与最终的用户沟通环节,强调了清晰的故事线在数据价值实现中的决定性作用。对我而言,这本书最大的贡献在于,它帮助我建立了一种更具韧性的思维模式,使我能够从容应对数据量级爆炸和技术迭代加速带来的冲击。它不是教我如何处理特定的“非结构化数据”,而是教会我如何与“结构未定”的世界打交道。这种更高维度的认知迁移能力,才是真正体现了这本书的超凡价值,让我在面对任何新的、陌生的信息挑战时,都能迅速找到切入点,并构建起有效的分析路径。
这本书的封面设计充满了现代感,简约的配色和大胆的字体选择立刻吸引了我的目光。初读前几页,我立刻被作者对于数据价值的独到见解所折服。他不仅仅是将数据视为冰冷的数字,而是将其描绘成隐藏在日常噪音中的宝藏,需要独特的“工具”去发掘。书中对传统数据处理范式的批判性审视,让我深思我们当前许多分析工作的局限性。特别是关于如何构建一个能够适应快速变化的数据环境的思维框架,这部分内容极为深刻,它提供了一种超越工具本身的哲学高度,去理解信息时代的本质。我尤其欣赏作者在论述中穿插的那些精心挑选的案例研究,它们并非是教科书式的枯燥陈述,而是生动地展示了理论如何落地生根,如何从看似混乱的源头提炼出清晰的商业洞察。阅读过程中,我感觉自己像是在跟随一位经验丰富的向导,探索一片广阔而未知的领域,每走一步都有新的发现和启发。这种感觉是难以言喻的,它混合了挑战的兴奋和解决难题后的满足感。这本书的结构安排得非常巧妙,知识点的递进逻辑清晰流畅,即便是对于数据分析领域的新手,也能平稳过渡到更高深的理解层面,这得益于作者对复杂概念的精妙拆解和生动类比。
拿起这本书的时候,我本来预期会读到一堆关于特定算法或软件操作的手册式内容,但事实远比我想象的要精彩得多。这本书的重点似乎完全不在于“如何操作”,而在于“如何思考”。作者花费了大量的篇幅来探讨数据“质地”对分析策略的影响,这一点在国内的很多同类书籍中是很少被深入探讨的。他非常坦率地讨论了在处理那些边界模糊、缺乏明确标签的数据流时,人类直觉和机器智能之间那种微妙的、相互依存的关系。其中有一章详细分析了文本情感分析中的“语境漂移”现象,举例非常精彩,让我对过去一些失败的项目有了豁然开朗的认识。这种对分析边界的探索,使得这本书更像是一部关于“数据侦探学”的指南,而不是一本技术手册。我甚至觉得,这本书的价值更偏向于管理层和战略制定者,因为它迫使读者跳出技术细节的泥潭,去思考如何建立一个能够容忍和拥抱不确定性的组织文化。我反复阅读了关于“噪音过滤与信号增强”的章节,作者提出的那些非线性模型假设,彻底颠覆了我过去对数据清洗的刻板印象。