Data Management Glossary
A reference guide of key data management terms and definitions, available exclusively to DAMA members. 2,578 terms and definitions. Use the A–Z sidebar or the search bar to find a term.
To terminate a processing activity before it finishes, typically because one of the following occurred: an invalid input, a system error, or a lack of resources.
Less specific in representation, or without a relation to a specific instance. Does not mean “more generalized.” See also generalization.
In data modeling, the redefinition of data entities, attributes, and relationships by removing details to broaden the applicability of data to a wider class of situations, often by implementing supertypes rather than subtypes.
In data services, the process of layering virtualization between data and its source. It redefines the data attributes or relationships by hiding details of the location, entities, and/or relationships of the information to broaden the applicability of data structures to a wider class of situations (i.e., implementing supertypes rather than subtypes, data access objects, data services, etc.).
The process of partitioning a model into smaller subparts for presentation. Used in data modeling to show related areas in a more readable scale.
The presentation of all or part of a model detail. Used in data modeling to show higher levels of entities and relationships to illustrate the basic subject area contents.
Generally, the ability to obtain or make use of something.
In data management, the operation of reading or writing information.
To obtain or retrieve.
The ability to readily obtain data when needed.
Degree of correctness or conformity to a true value.
Level of incidence of error; accurate records are error free and can be used as a reliable source of information.
In data management, how closely a piece of data reflects its true, real-world value. Data accuracy is the first and most critical component of a data quality framework.
A machine learning approach where the algorithm selects the most informative data points to learn from.
A contribution to the performance of a function or process. An activity is at a lower level than a function or process, but at a higher level than a task or step. Inputs, activities, and outputs combine to form a process.
One of the DAMA-DMBOK Functional Framework Environmental Elements. Each function is composed of lower-level activities, which may be grouped into sub-activities, and then further decomposed into tasks and steps.
The role of a person or an external system or device that interacts with the system, business, or component under analysis.
In general, not cyclic; not composed of regular cycles.
A characteristic of a graph where there exists at most one path between any two nodes. See also connected.
A query constructed and executed to answer an immediate and unanticipated question or need, in contrast to a planned query. For example, a dynamic SQL SELECT statement against a relational database, constructed by a knowledge worker using an English-like or point-and-click interface of a desktop-resident business intelligence tool. The data returned may drive further analysis and reporting. Contrast with planned query.
Systems that change their behavior based on feedback.
The world’s first operational packet-switching network, developed by the US Department of Defense. Precursor to the internet, which evolved into the World Wide Web.
An input designed to fool a machine learning model into making incorrect predictions.
An analysis technique that relates occurrences of activities by individuals or groups. Market basket analysis is a type of affinity analysis.
An autonomous entity that perceives its environment and acts upon it to achieve goals.
Data resulting from processes that combine and summarize atomic data.
Generally, the process of gathering parts into a whole.
In data management, a process that transforms atomic data into aggregate-level information by using an aggregation function such as count, sum, average, standard deviation, etc.
An iterative project management framework that breaks projects down into several phases, commonly known as sprints.
A group of software development methodologies based on iterative and incremental development, where requirements and solutions evolve through collaboration between self-organizing, cross-functional teams. See Agile methodology.
A set of rules or steps that will result in a defined end from a defined start.
Generally, an alternative reference to a standard name or term.
In relational database management systems, a database object that indirectly references another database object; for example, an abbreviated table reference within an SQL query.
In a distributed environment, an object that refers to another object to avoid having to use the full location qualifier of the other object. This alias is not dropped if the object referred to is dropped.
The first version of something released to a formal testing team.
A primary key that is valid and acceptable, but is not the preferred primary key. See also key, alternate.
Uncertainty in meaning or reference, depending on the context or usage. An ambiguous reference may have multiple meanings in the absence of context or usage specifications.
In the United States, a large continuous demographic survey that is sent to residents on a monthly basis, rather than decennially. It contains more demographic questions than the old census long form and provides more up-to-date information than was previously collected.
A private not-for-profit organization that coordinates the development and use of voluntary consensus standards in the United States and represents the needs of US stakeholders in worldwide standardization forums. Formerly the American Standards Association, from which we get the American Standard Code for Information Interchange (ASCII).
A signal represented by an oscillating wave rather than digital pulses.
The process of systematically collecting, cleaning, transforming, describing, modeling, and interpreting data.
A person who performs analysis or is skilled in analysis. See also business analyst; business systems analyst; data analyst; systems analyst.
Software that packages business intelligence technology to support a specific knowledge-driven business process.
The system of criteria and standards within which data are analyzed.
Business intelligence procedures and techniques for exploration and analysis of data to discover and identify meaningful information and trends.
A software tool that can process and analyze large quantities of data to aid decision-making.
In the context of data, points in the data that are significantly different from the others and do not conform to expected patterns or norms. Anomalies can manifest as errors, inconsistencies, or unexpected patterns, often impacting data quality, analysis, and decision-making.
Data that originally contained personally identifiable information (PII) but has been processed to remove any indications that could positively identify an individual person.
The standard form of SQL concurrently defined by the American National Standards Institute (ANSI) and International Organization for Standardization (ISO) and first released in 1986. Provides a consistent, vendor-neutral framework for relational database management systems, ensuring interoperability across such platforms as MySQL, PostgreSQL, SQL Server, and Oracle.
Relevance to the current subject.
Ability to be put to specific use.
In computing, software functions and services implemented together to support one or more related business processes.
The process of building and maintaining software applications.
The IT organization responsible for application development. Synonymous with software development or software engineering.
A published standard format for communicating with applications.
In a three-tier application architecture, the middle tier of software (and possibly hardware) where business logic is performed.
A company offering network access to application programs and services for other parties. ASPs typically provide the applications, infrastructure, and technical support for a monthly service charge.
A programming technique where each module is independent: it has no dependency on, is unrelated to, and does not communicate with all other modules.
In graph theory, a connection between two nodes in a graph. Also known as an edge.
In mathematics, a curved line.
Generally, a person trained in the planning, design, and oversight of the construction of something, usually buildings.
In information technology, an experienced and skilled designer responsible for architecture supporting a broad scope of requirements over time, beyond the scope of a single project. The term implies a higher level of professional experience and expertise than an analyst, designer, modeler, or developer.
A way of thinking about and understanding architecture and the structures or systems requiring architecture.
Generally, the design of any complex object or system, including the implied architecture of abstract things such as music or mathematics, the apparent architecture of natural things such as geological formations or living things, or explicitly planned architecture of human-made things such as buildings, machines, organizations, processes, software, and databases.
In data management, the organized arrangement of components to optimize the function, performance, feasibility, cost, and/or aesthetics of an overall structure.
In common use, the art and discipline of designing buildings and structures, from the macro level of urban planning to the micro level of creating furniture and machine parts.
A set of standard programming structures, design patterns, formats, and protocols for how software applications should operate and communicate with each other.
A master blueprint for an organization’s existing and planned portfolio of software applications, how they support the organization’s processes, and how they interface with each other and with the organization’s databases.
The portion of an enterprise architecture that describes organizational goals, roles, reporting structures, and locations, but excluding the enterprise data architecture, process architecture, technology architecture, and application architecture.
The overall design and implementation of components of the business intelligence environment, including:
- the data warehouse, data marts, and staging area databases;
- the flow of data from operational sources and these databases;
- the selection and configuration of database management systems and database administration tools used for business intelligence;
- the selection and configuration of data integration products used to extract, cleanse, transform, and load data;
- the design patterns and standards for data integration programs;
- the selection and configuration of business intelligence software products that enable access, reporting, and analysis;
- the data schemas presented to business professionals for ad hoc query and analysis;
- the user interfaces for query, analysis, and reporting; and
- the administrative controls put in place to safeguard the data.
The future-state business process models of an enterprise, used in conjunction with a business data architecture to perform information value chain analysis. Part of an enterprise architecture.
A distributed technology approach where application software processing is divided by function. Servers perform shared functions such as processing business rules, managing communications, managing databases, or providing print services. Clients performs individual user functions – providing customized interfaces, performing screen to screen navigation, offering help functions, etc. Client and server software may reside on the same hardware platform, but each component is designed to be distributed across a networked environment for efficiency.
An architecture where only the original manufacturer can make add-ons and peripherals.
In common usage, the physical technology infrastructure supporting data management, including database servers, data replication tools, and middleware.
The method of design and construction of an integrated data resource that is business driven, based on real-world subjects as perceived by the organization, and implemented into appropriate operating environments.
Generally, an integrated collection of models and design approaches used to align information, processes, projects, and technology with the goals of the enterprise. These high-level design artifacts typically describe target views of the enterprise. Enterprise architecture may include:
- an enterprise data model,
- related data integration architecture,
- a business process model,
- an application portfolio architecture,
- an application component architecture,
- an application component architecture,
- an organizational business architecture, and
- the enterprise information value chain analysis that identifies the linkage and alignment across these perspectives, and to enterprise goals.
Other models and other forms of architecture may also be included within the enterprise architecture.
In the Zachman Framework for Information Systems Architecture, the enterprise architecture generally includes design artifacts identified in Rows 1 and 2 (conceptual views of data, process, locations, events, roles and goals), the value chain analysis describing the linkages between these perspectives, and high-level decisions about how to implement technology supporting these concepts in an integrated manner.
The analysis and design of the data stored by information systems, concentrating on entities, their attributes, and their relationships.
The integrated set of design artifacts defining how data (including the logical data model), applications (including the application portfolio architecture and data integration architecture), and technology (including portfolios of technology products and standards) will integrate to support the business architecture.
The design for integration of metadata across data dictionaries, directories, and repositories.
A form of architecture where the user interface layers, the application processing layers, and the data management layers are all logically separate parts which communicate through services. Also called n-tier architecture. See also architecture, three-tier.
The published specifications for a computer by a vendor, allowing other companies to create add-ons to enhance and customize the machine, and to make peripheral devices that work properly with it. In practice, it has been difficult to engage on a corporate basis due to the risk involved in a source that has multiple editors and has little to no assurance of quality when in use. Outsourcing the risk to a second party, who then uses the open source and accepts the liability for the code, is then the way to engage with open-source code.
The structural design of process systems, such as computers, businesses, or other complex systems.
Enterprise process architecture typically includes
- a functional decomposition,
- process flow diagrams, and
- value chain analysis linking processes to data (subject areas or entities), organizations, roles, goals, applications, and/or projects.
Part of a technology architecture, identifying selected vendor-specific software tools and services. Although not implied in the name, it may also include industry-wide standards and protocols.
A representation of a system, including a mapping of functionality onto hardware and software components, and the human interaction with these components. Includes applications, software components, interfaces, and projects.
The master plan for the information technology technical infrastructure depicted in diagrams and specifications of hardware and system software products, locations, configurations, standards and adopted protocols, along with linkages of computing platforms and/or servers to existing and planned applications and databases.
A structure for a database environment consisting of a presentation tier, an application tier, and a data tier. The presentation tier is seen and used by the programmers and other users of a DBMS, also called the user schema or external schema. Presentation tiers can overlap. The application tier is the combination of all the defined structures in the presentation tier for a given database, also called the logical tier, data access tier, or middle tier. There may be additional data in the application tier that is not in any presentation tier. The data tier is the database administrator’s view of the database, also called the internal schema. The data tier is the definition of the physical storage structure of a database.
A copy of a database or documents preserved in a secondary, lower-cost storage location, for infrequent historical reference and/or recovery.
To move stored data (structured or unstructured) to a secondary, less readily accessed location, at lower storage costs, for historical reference and/or recovery.
The number of object roles that a single predicate (relationship type) can have in its diagram representation. See also n-ary.
A grouping of similar items of the same storage type in a sequential pattern and referenced by a sequential index value. See also matrix.
A tangible output from an activity or task. For example, a logical data model, a requirement document, or a project plan.
Software that performs a function previously ascribed only to human beings. It can encompass such things as natural language processing, simple automated tasks, and complex problem-solving.
A computational model inspired by the human brain, used in machine learning.
A form of artificial intelligence (AI) that surpasses human intelligence in all aspects, including creativity and problem-solving. ASI represents intelligence far beyond human capabilities, such as autonomous self-improvement and the execution of any intellectual task more effectively than humans.
A common code used in transmitting information over networks, using 7 or 8 data bits, a parity bit, and a stop bit. See also EBCDIC (Extended Binary Coded Decimal Interchange Code).
Generally, something that has value or produces benefit.
In accounting, something of value on a balance sheet.
Describes how an asset or a service will perform in objective and measurable terms. The measurement is sometimes as simple as assigning a number. An example would be a range of 1 to 5, where 1 = poor and 5 = excellent.
Nonphysical asset, such as accounts receivable.
Physical asset, such as equipment.
To determine relationships between entities, including characteristics of the relationship: dependent or not (optional, orphan), exclusive (at most one) or not (multiple). See also relationship.
See relationship.
In statistics, any relationship between measured quantities that shows a statistic dependency.
In object-oriented programming, a relationship between object classes that enables an object instance to perform an action on another’s behalf.
The largest and oldest international scientific and educational computer society.
A style of communication in which the initiator does not wait for a reply. Opposite of synchronous.
Data replication where the target database is updated as soon as possible after updates occur to the source database, but not as part of a single integrated transaction. Failure to update the target has no impact on the source database. Sometimes referred to as near-real-time replication.
Data at the lowest chosen level of detail (granularity). The level of detail chosen depends on the information requirements of the enterprise. For example, address could be one atomic item, or address could be split into further composite items, such as house identifier and city. Opposite of aggregate data.
Non-aggregated observations, or measurements of characteristics of individual units, which cannot be further decomposed and retain any useful meaning.
Standard properties of relational databases.
A mechanism in neural networks that focuses on specific parts of the input data when making predictions.
A specific characteristic or property of a data entity that helps to define its nature and behavior within a dataset. In the context of data analysis and data science, attributes are crucial as they provide the necessary information to understand and manipulate data effectively.
A formal and official verification of validity, accuracy, and conformance to requirements, regulations, standards, and/or guidelines.
Data maintained to trace activity, such as a transaction log, for purposes of recovery or audit.
The process of adding to something to make it more or greater than the original.
In logic, a relationship where if X leads to Y, then XZ will lead to YZ.
In data security, the process of verifying whether a person or software agent requesting a resource has the authority or permission to access that resource.
In data quality, the process of verifying data as complying with what the data represents.
A source of data or information that is recognized by members of a community of interest (COI) to be valid or trusted because its provenance is considered highly reliable or accurate. During the life cycle process, the authoritative source (or system of use in which it is housed) can evolve according to use. Subject matter experts validate that the data is authoritative, and data management assures that data from the authoritative source is provided to users, and that it is current.
In data security, the granting of authority allowing a person, group, or software agent to access a resource.
In data security, a request to grant authority to a person, group, or software agent to access data for which the data consumer does not have access privileges.
A method of automatically identifying and collecting data on items and then storing the data in a computer system. For example, a scanner might collect data about a product via an RFID chip.
The act of replacing control of a manual process with computer or electronic controls.
The percentage of time a system or data resource is accessible or expected to be accessible for productive work.
A figure that shows data using network or relational models. Named after Charles Bachman. Also called a data structure diagram.
To take a copy of a system to ensure its continued availability in the event of a hardware or software failure requiring recovery of the database to restore the data.
backup: The copy of the system information and data used for recoverability.
In computer architecture, the server-side part of a system that handles logic, data, and services behind the user interface.
In data warehousing, the 80 percent data management activities of enterprise data warehousing, such as data profiling, data modeling, and extract-transform-load processing.)
An Agile term for a list of desired user requirements, some of which may be a wish list. It identifies which features to implement first.
An AI core algorithm in training neural networks, where the error is propagated backward to adjust weights.
A backup snapshot taken while the system is offline.
A backup snapshot taken while the system is online.
Able to accept input from older or earlier versions of a device or software.
Operational on older technology, even if limited in functionality.
A strategic performance-management tool consisting of a semi-standard structured report, supported by proven design methods and automation tools, that can be used by managers to keep track of the execution of activities by staff within their control and monitor the consequences arising from these actions. It provides a comprehensive, top-down view of organizational performance measurements with a strong focus on vision and strategy, based on concepts developed by Robert Kaplan and David Norton.
Criteria used to evaluate the qualification of companies for the Malcolm Baldrige National Quality Award; leading management practices used to measure organizational performance in seven categories: Leadership; Strategy; Customers; Measurement, Analysis, and Knowledge Management; Workforce; Operations; and Results.
In data warehousing, the normalized data structures maintained in a data warehouse, in contrast to the denormalized dependent data mart tables sourced from the base tables.
Outside of data warehousing, a table for an entity that is not dependent on any other entity in the database.
The unit used as the basis of an index number, or to which a constant series refers. For example: base period, base weight, base currency.
International banking supervision standards designed to ensure the liquidity of financial institutions doing business in European Union countries.
A method of processing high volumes of data where a group (or batch) of data is collected over a period of time and then processed together.
What something does at any point in time. The execution or carrying out of a process constitutes behavior. Behavior is something that happens, as opposed to something that is. Opposite of state.
Confidence in inherent truthfulness.
A statistical frequency distribution pattern that is shaped like a bell (narrow at ends, wide in the middle of the range). See also normal distribution.
A point of reference for measurement, comparison, and evaluation. A benchmark can be a standard of excellence or a point-in-time snapshot measurement for comparison with other benchmarks. A benchmark may be an internal or external measurement.
To analyze and compare an organization’s processes (an internal benchmark) against the performance of another organization or an industry standard (an external benchmark).
The process of comparing the performance of artificial intelligence models using standard datasets and evaluation metrics.
A technique, method, process, discipline, incentive, or reward generally considered more effective at delivering a particular outcome than other means.
A release of software to a limited population, under controlled conditions, to test for functionality completeness and execution correctness.
Generally, a distortion of something to support a particular view.
In data analysis, a distortion of data or information that affects its interpretation, or a distortion of interpretation that supports a particular view.
In machine learning, a model’s error due to incorrect assumptions in the learning algorithm.
A distortion of fact interpretation based on sole use of data provided by or pre-selected by the sponsor of the research, which may be skewed toward a certain result, rather than being completely objective.
A distortion of fact interpretation due to non-random selection of sample contents.
A distortion of fact interpretation based on sole use of data that supports the desired outcome, rather than a complete dataset.
A distortion of fact interpretation by only using the results that support the desired outcome, and ignoring or not displaying the other results.
A naming convention for binary relationships where the relationship is described twice, in sentences, once with one entity named as the subject and the other entity as the object of the sentence, and again with the first entity as object and the second as subject.
Data volumes that are exceptionally large, normally greater than 100 terabytes and more commonly in the petabyte and exabyte range. Big data is used in data warehousing and analytic solutions where the volume of data poses specific challenges unique to very large volumes of data, including data loading, data modeling, data cleansing, and analytics.
Generally, an exchange of something between a sending organization and a receiving organization in which all aspects of the exchange process are agreed between counterparties.
In data management, an exchange of data and/or metadata between a sending organization and a receiving organization in which all aspects of the exchange process are agreed between counterparties, including the mechanism for exchange of data and metadata, the formats, the frequency or schedule, and the mode used for communications regarding the exchange.
The information and relationships that document the entire lifecycle of a product. Includes the associated product information (administrative, programmatic, technical, and financial) and its location.
A list of raw materials, down to the atomic level necessary to create a final item.
Consisting of two components or values.
The format of a compiled and linked program that is ready to execute on a specified system.
A unit of measurement for data based on the binary number system using zero and one. (From binary + digit.)
A subset of personally identifiable data consisting of a person’s physical or physiological attributes, such as a retinal scan, fingerprint, or DNA.
The amount of bits transferred over a conduit or connection in one second.
A decentralized, tamper-resistant digital ledger of all transactions across a distributed network, which can be used to record transactions across many computers so that any involved record cannot be altered without the alteration of all subsequent blocks.
The situation in which one process locks a resource that another resource needs. The second resource is “blocked.”
A type of website containing regular entries of commentary, notes, or links to graphics or video. Short for “web log.”
The sum of all professional knowledge in a given field, or what is generally accepted to be true.
A body-of-knowledge document available from the International Institute of Business Analysis.
A body-of-knowledge document created by the Association of Business Process Management Professionals International.
A body-of-knowledge document created by the business technology architects’ association Iasa Global (formerly International Association of Software Architects).
A project undertaken by the Canadian Information Processing Society to outline the knowledge required of a Canadian Information Technology Professional.
A guide to knowledge about data management published by DAMA International.
An archival body-of-knowledge document created by the International Association of Software Architects. Replaced by Iasa Global’s Business Technology Architecture Body of Knowledge (BTABoK).
Published by the International Information System Security Certification Consortium (ISC2). Contains the information tested to achieve the Certified Information Systems Security Professional designation.
Developed by Carnegie Mellon University, a body of knowledge on personal software process and team software process.
An acronym and registered trademark for the Guide to the Project Management Body of Knowledge, a publication of the nonprofit Project Management Institute (PMI) and an internationally recognized standard (IEEE Std 1490-2003) defining the fundamental vocabulary of project management and identifying generally accepted project management practices.
The Guide to the Software Engineering Body of Knowledge, a book published by the Institute of Electrical and Electronics Engineers (IEEE) Computer Society.
A marker used to save a place in a book or dataset, or an internet address.
Relating to or of an algorithm or calculation that results in only a True or False result. Named for George Boole.
Logical operators that combine propositions to evaluate to only a True or False result. Includes and, or, if–then, except, and not.
A search method using Boolean operators (and, or, not) to focus the search.
Also called Boston box. See chart, portfolio.
Short for robot. A computer program that runs on the internet and performs automated tasks for other users or programs. Bots can imitate human behavior, but they are faster and more accurate than humans at performing repetitive tasks. Some bots are designed with malicious intent and can have a negative impact on websites or applications.
A malicious bot. There are several kinds of botnets, including scam bots that harvest email addresses; website scrapers that grab content from websites and use it without permission; and spy bots that collect data about users without their permission.
Standards for defining process flows using web services.
Standards for defining process flows controlled by web services.
In databases, a software function that prevents users from querying a database once transaction loads reach a certain level.
A software mechanism to move large data files that uses compression, blocking, and buffering to optimize transfer times.
In data warehousing, a tabular representation of the intersection of shared dimension tables with data subject areas, data processes, data facts, data marts, etc.
Generally, any purposeful activity.
Specifically, a commercial or industrial enterprise. Commercial activity engaged in as a means of livelihood.
A set of methods or procedures that may be executed in the form of transactions relative to a business. See also activity; business process.
The ability to automatically monitor events in an executing business process through immediate notification, thanks to a sophisticated technical infrastructure.
The study of business processes, practices, and business systems requirements.
The application of information to better understand business opportunities and challenges. See also business intelligence (BI).
Generally, a knowledge worker responsible for interpreting data, performing calculations, and distributing reports to other knowledge workers.
In data management, a professional responsible for understanding the business processes and the information needs of an organization, serving as a liaison between IT and business units, and acting as a facilitator of organizational and cultural change. See also business systems analyst; systems analyst.
Metadata that includes data definitions, report definitions, users, usage statistics, and performance statistics.
A structured format for organizing the reasons, benefits, and estimated costs for initiating a project or program.
The degree of uninterrupted stability of an organization’s systems and operations in spite of potentially disruptive events.
Data about people, places, things, rules, events, or concepts used to operate and manage any enterprise (not just commercial enterprises). Used to identify data that is not considered to be metadata.
A knowledge worker, business leader, and recognized subject matter expert assigned accountability for the data specifications and data quality of specifically assigned business entities, subject areas, or databases, but with less responsibility for data governance than a coordinating data steward or an executive data steward.
Data created while operational processes are in progress. Examples include customer orders, cash withdrawals, and supplier invoices. Also called transactional business data.
Connects data governance activity with business needs. Also called data governance strategy map.
A course studying design, import, and manipulation of text, graphics, audio, and video used in presentation management and publishing systems.
An organization’s continuously increasing, constantly changing need for current, accurate, integrated information, often on short notice or very short notice, to support its business activities. It is a very dynamic demand for information to support the business that constantly changes.
A set of concepts, methods, and processes to improve business decision-making using any information, from multiple sources, that could affect the business, and applying experiences and assumptions to deliver accurate perspectives of business dynamics.
Data that helps a data warehouse administrator manage a data warehouse, such as user profiles and data access history.
An IT professional specializing in assisting and supporting business professionals in becoming more self-sufficient in the use of query, reporting, and analysis procedures and tools. A business intelligence (BI) analyst trains knowledge workers, assists them in solving more complex analytical and reporting problems, and provides second-level support for user problems with BI data and tools, and may serve as the administrator for the BI environment.
An IT professional with overall responsibility for the business intelligence environment, its architecture, and the effectiveness of knowledge workers engaged in business intelligence (BI). A lead BI analyst. May also be the data warehouse architect, or these responsibilities may be distinct.
An IT professional software developer who specializes in report writing and/or the development of analytic applications.
The hardware, software, and organizational support for business intelligence activity that enables knowledge workers to access, analyze, and manipulate data. It generally includes the business intelligence software, user interfaces, associated infrastructure hardware and software, data mart databases, and multidimensional data cubes. It may also include data warehouses and the data integration programs that provide data for business intelligence.
The infrastructure of selected enabling tools and technologies necessary for the development and deployment of business intelligence applications.
An application service provider providing data warehousing and business intelligence capabilities as outsourced services hosted off-site. A BISP ties into source information systems and databases behind a corporation’s firewall, providing traditional data warehouse and analytic application capabilities to internal knowledge workers and external customers. Often used to extend business intelligence functions into e-commerce.
The creation, publishing, and sharing of custom business analytics reports and dashboards by end users of cloud technologies.
Technology and products (tools) used by knowledge workers to access data, analyze, and share information; understand business performance; and improve decision-making. Includes query and reporting tools, online analytical processing (OLAP) technologies, statistical analysis tools, data mining tools, scenario modeling tools, planning and budgeting tools, advanced analytic applications, dashboards and scorecards for performance monitoring, and enterprise reporting tools.
The training and assistance available to business professionals in use of business intelligence tools and techniques and the valid interpretation of BI data. Typically, a help desk provides Level 1 support, with BI analysts providing Level 2 support.
A current or future-state representation of some aspect of an enterprise, typically from a process, data, geographic, event, organizational, or financial perspective.
A container for application data, such as a customer or an invoice. Data is exchanged between components by business objects.
An umbrella term for the methods, metrics, processes, and systems used to monitor and manage the performance of any enterprise.
The use of techniques and tools to measure performance against specific key performance indicators, often coupled with comparative information from industry sources. Dashboards support business performance measurement. The Balanced Scorecard (BSC) is a specialized form of business performance measurement.
The use of techniques and tools to understand the factors affecting business performance and to explore “what if” scenarios to help consider the implications of alternative courses of action. See also scenario modeling.
A collection of related and organized functions, activities, procedures, steps, or tasks that follow a defined sequence to produce a product or service addressing a particular business requirement or objective.
Activities that may involve multiple stakeholders, including employees, technology, and tools, and are designed to ensure efficiency, consistency, and measurable outcomes.
The design, monitoring, and control of complex interactions between people, applications, and technologies designed to create customer value.
A model that defines maturity levels for business processing from Initial, through Managed, Standardized, and Predictable, to Innovating. Developed by the Object Management Group.
A model of the functions, activities, and procedures performed in any organization. It may consist of:
- A context diagram showing the relationship of the overall process to those outside the model’s scope, along with the inputs to and outputs from the overall process;
- One or more functional decomposition diagrams showing how the overall process is made up of contributing processes at lower levels (a “vertical view”);
- One or more process flow diagrams showing how the outputs of one process serve as the inputs to other process (a “horizontal view”). The process flow may be cross-functional or within a single function;
- One or more business process model diagrams, each depicting the inputs, outputs, start and end events, component activities, roles, and metrics of a single process;
- The business definition of each process; and
- The value chain analysis of the process, identifying relationships to data, organizations, roles, and systems.
A stylized approach to graphically documenting the definition, objectives, start and end events, inputs, outputs, component activities, roles, and metrics of a single process. Sometimes referred to as a context diagram, but a BPM diagram includes more information than a traditional context diagram.
The standard for business processes diagrams. It is intended to be used directly by the stakeholders who design, manage, and realize business processes, but at the same time be precise enough to allow BPMN diagrams to be translated into software process components. Maintained by the Object Management Group.
A form of outsourcing that involves transferring responsibilities for entire specific business functions or processes to a third-party provider.
The process of analyzing and radically transforming existing business activities, eliminating or minimizing costs and maximizing value in order to achieve breakthrough levels of performance improvement.
A professional responsible for understanding the business processes and information needs of an organization, serving as a liaison between IT and business units, and acting as a facilitator for organizational and cultural change. See also systems analyst.
A method for defining an enterprise architecture and information systems architecture developed by International Business Machines Corporation (IBM) in the early 1980s.
An event involving the exchange of products, money, and/or information.
Commerce transactions between equivalent businesses, such as between a wholesaler and a retailer.
Commerce transactions between a business and a consumer, such as in a retail sale.
Commerce transactions between a business and a governmental body, such as between a business and an elected water commission.
A single character of data stored electronically in 16 binary bits. A datum.
The term originally coined by International Business Machines Corporation (IBM) with the announcement of the 360 series of computers in 1974. Originally consisted of 8 bits; could be used to store a single character, digit, or two decimal digits (“packed decimal”), or in combination could be used to store numbers. ASCII and EBCDIC are the two dominant character coding schemes based on 8 bits.
A block of memory on a computing system for temporary storage of data likely to be used again.
A state when a data request can be supplied from data within a cache, rather than directly from disk storage.
A state statute to enhance privacy rights and consumer protection for residents of California.
The part of an organization that handles inbound/outbound telephone or email communications with internal and/or external customers. An information technology (IT) help desk is a call center for customers of the IT department.
Detailed tracking, reporting, and analysis that provides precise measurements regarding current marketing campaign efforts, their performance, and the types of leads they attract.
A set of one or more columns that can uniquely identify each row in a table, providing that no subset of its columns can also uniquely identify the row. A table can have multiple candidate keys but only one is chosen as the primary key. The candidate key’s attributes can contain a NULL value that opposes the primary key.
An accepted principle or role; a body of principles, rules, standards, or norms by which something is judged.
A data model of the inherent structure of data without regard to applications, hardware, or software implementations. Built according to specific canons. Usually a result of canonical synthesis.
The concept that if everyone followed the canons (rules) for developing a data model, then those independent data models could be readily plugged together, just like a picture puzzle, to provide a single, comprehensive, organization-wide data architecture.
A system development process capability maturity model published by the Software Engineering Institute at Carnegie Mellon University. The CMM is a guide to improve an organization’s software development process, featuring defined practices used to rank an organization at one of five process maturity levels. Preceded by the Capability Maturity Model for Software. Commonly called Capability Maturity Model (CMM).
The maximum amount that can be held, contained, or processed at one time.
A number measured on a scale with an arithmetically meaningful zero point. Generally used to measure quantities or volumes. Can be manipulated by all the binary operators: exponentiation, multiplication and division, addition and subtraction, comparison (e.g., less than), matching, and Boolean. See also ordinal number; interval number; nominal number.
The number of entities or members in a set that may participate in a given relationship, expressed as one-to-one (1:1), one-to-many (1:M), or many-to-many (M:N).
A set of points on a set of axes used to show location or proximity.
In data processing, given two or more populations, the set of all possible combinations, taking one value from each population. Also called Cartesian join. Usually a large and meaningless answer set for an incorrectly phrased query.
The study and practice of making maps. Maps function as visualization tools for spatial data. Most quality maps are now made with geographic information system (GIS) software and databases.
The declaration made on a hierarchical (one-to-many) relationship between parent and child, that a request to delete a parent instance will also result in deleting the related child instances. Usually associated with a foreign key (which defines a hierarchical relationship), with the referring entity table (where the foreign key is stored) being the child and the referenced entity table is the parent.
An evaluation of an instance of a process to determine what environmental or inherent attributes drove success or failure of the process.
Generally, a complete list of things, usually arranged systematically.
In databases, the component of a database management system (DBMS) where metadata about DBMS objects is stored. Most relational DBMS products keep the catalog as relational tables. The majority of metadata in a DBMS catalog is technical metadata (e.g., names, types, lengths, occurrences, keys) collected automatically by the DBMS software, although business definitions can be added as comments.
A (data) catalog is an active data dictionary.
The generic term for items at any level within a classification.
A scheme made up of a hierarchy of categories, which may include any type of useful classification for the organization of something.
The relationship between cause and effect, which is important for understanding the impact of actions in artificial intelligence systems.
Generally, any small compartment.
In multidimensional design, a data point defined by one member of each dimension of a multidimensional structure. Often cells in multidimensional structures are empty, leading to “sparse” storage.
A team of people that promotes collaboration and use of best practices around a specific focus area to drive business results.
A centralized data management services organization of data management professionals.
The part of a computer that reads, interprets, and performs instructions.
A token of authorization or authentication.
In data security, a computer data security object that includes identity information, validity specification, and a key.
The process of reviewing something to verify it meets established standards.
A professional certification offered by TDWI (Transforming Data with Intelligence), using examinations developed and delivered by the Institute for Certification of Computing Professionals (ICCP).
Data that has passed data quality review, certifying it meets established standards.
A professional certification program offered by DAMA International, using examinations developed by DAMA International.
The documentation of ownership of something, from capture, through possession, storage, and management, to disposition. This is especially important for compliance documentation. See also data provenance.
Connecting a series of commands or responses.
In cryptography, a method of encryption where each block defines or contributes to the encryption of the following blocks.
The process of coordinating changes to a system in order to minimize change-related errors and therefore improve data quality and system availability. Proposed changes are reviewed and evaluated for related impacts, grouped and scheduled, and implemented and migrated through various test environments before being implemented into the production environment. Database change control disciplines are a very important responsibility of database administrators.
The process of capturing changes made to a production data source. Change data capture is typically performed by reading the log file of the database management system of the source database. Change data capture consolidates units of work, ensures data is synchronized with the original source, and reduces data volume in a data warehousing environment.
A distinguishing feature or quality.
Pertaining to, constituting, or indicating the character or peculiar quality of a person or thing; typical; distinctive.
An abstraction of a property of an object or of a set of objects.
A visual representation of data, using shapes, colors, symbols, graphs, images, tables, diagrams, etc. to show patterns, relationships, or ideas, that makes the data easier to understand or gives context to create some form of information.
A form of visualization that shows patterns of ideas or data by grouping them by topic or some attribute they share.
A chart showing multiple lines from left to right, each of which defines the top line of an area within the chart. The areas are marked with colors, textures, and/or hatching. The areas may be overlapping or stacked.
A chart that uses a geographic map of the world with the size of countries or their subdivisions distorted by the value of a property of that area, such as population.
A chart that shows bars to illustrate frequencies or values for individual categories.
A chart that displays five values of a measurement where the median and quartile values are the center and edges of a box, and the lowest and highest values are the ends of lines extending from the box.
A diagram illustrating steps necessary for two disparate positions to come to a consensus in a middle area, crossing some gap between them, usually illustrated by a bridge over a river or chasm.
A chart showing two dimensions on horizontal and vertical axes, and a third dimension in the size of the points.
A variation of the bar chart that compares a single, primary measure to one or more other measures, such as a target or a quantitative scale, displayed in qualitative ranges (poor, fair, good, etc.) by using variations of hue for a single color (which is helpful for colorblind vision). These long narrow graphs can be grouped to save space, especially on web forms or dashboards.
A chart showing bars representing range of value change within a point’s time interval.
A chart consisting of a geographic map modified to show some measurement of the map’s area, contents, or qualities. Modifications can be to color or to proportional size. There are two types: area cartograms and distance cartograms.
A chart with the X-axis showing a unit of measure and the Y-axis showing a rate per unit. Boxes show the result of X units × Y rate for a specific segment, such as customer. Tall thin boxes above the X-axis are desirable; long short boxes above the X-axis are less desirable; boxes below the X-axis are undesirable.
A visual representation of a system’s feedback loops, where positive loops cycle clockwise and negative loops cycle counterclockwise.
A chart that links an outcome to chains of possible contributing factors as a tree structure, working backward from an event to determine possible root causes, drawn sideways so that it resembles the skeleton of a fish. Because the chart resembles the skeleton of a fish, it is often called a fishbone diagram. A quality improvement concept invented by the Japanese statistician Karu Ishikawa.
A type of diagram that shows a system’s classes, contents, attributes, and relationships, including inheritance. UML is a common format for a class diagram.
A representation of objects and their links to each other, sometimes including time and/or sequence of relationships. Numbers show the sequence of activities or messages.
A visual representation of the parts of a whole, usually a system, with sequences of dependencies between components shown.
A form of visualization showing nested subsets of a set as circles within other circles, such as cities in states in countries in a total population, with each level represented as the area of one of the circles.
A form of visualization where a concept is decomposed into components to the right, “fanning” out levels.
A form of visualization showing relationships among concepts as arrows between labeled boxes, usually in a downward branching hierarchy. See also data model, conceptual (CDM).
A form of visualization that takes a tree diagram and turns it into a three-dimensional circle of attributes radiating from a parent.
A form of visualization for tracking process performance over time.
A form of visualization of the critical path for a set of interdependent activities, showing the longest discrete path through the tasks with the longest duration.
A form of visualization showing cycles of a concept’s stages, phases, or process steps in a clockwise path around a circle. Also called cycle diagram.
A form of visualization using a geographic map with overlaid data shapes using colors to illustrate ranges of values for each geographic block.
A graph of decisions and their possible consequences (including resource costs and risks) used to create a plan to reach a goal. Decision trees are a special form of tree structure constructed to help with making decisions. Regression trees approximate real-valued functions (e.g., estimate the price of a house or a patient’s length of stay in a hospital). Classification trees define the logic for categorization using Boolean variables such as gender (male or female) or game results (lose or win).
A chart consisting of a geographic map modified to show some relative travel times between points in a network.
A form of visualization showing one pool of two fixed resources shared by two entities. Each point shows a possible division of resources between the entities. Curves can be drawn between points of equal value to both parties according to the value associated with each resource. Named for creator Francis Ysidro Edgeworth.
A form of visualization that follows a process from a desired input through possible system events to final consequences. See also chart, fault tree.
A form of visualization showing top-down, deductive analytical steps through Boolean logic gates to all possible failure states. Also called failure tree.
A form of visualization where a topic appears in the center and forces for and against the topic are listed on each side.
An outline or hierarchy diagram depicting the hierarchical decomposition of processes into their component processes. A functional decomposition can be depicted vertically as an outline or horizontally as a hierarchy chart.
A form of visualization where inputs are drawn entering through the large end of a funnel and outputs are drawn leaving the small end.
A horizontal bar chart used in project management; a graphical illustration of a schedule that helps to plan, coordinate, and track specific tasks in a project. Named for Henry Gantt.
A form of visualization that divides the process of adoption of something into five cycles: trigger, peak, trough, slope, and plateau.
A form of visualization that uses a quartered chart comparing companies selling similar products according to their completeness of vision and ability to execute on that vision. The quarters are leaders (high on vision and execution), challengers (high on execution, low on vision), visionaries (high on vision, low on execution), and niche players (low on both vision and execution). Developed by Gartner, Inc., to evaluate vendors in specific market segments.
A chart where one set of values is represented by areas of rectangles and other sets of values are represented by colors. Used to look at large, fast-changing sets of structured data. In this chart, the size of a rectangle reflects its importance and color conveys the speed of change. Heat maps are often used in applications to monitor and analyze changes in stock market and portfolio data in financial services applications. Invented by Ben Shneiderman at the University of Maryland. See also chart, tree map.
A form of visualization showing positive and negative effects of some system or action by illustrating the positive items at the top as “heaven,” the negative at the bottom as “hell,” and neutral items in the center.
A form of visualization where a tree is displayed in a circular manner as a node-link diagram radiating out from the root rather than only descending. Also called hypertree.
A form of visualization with a medial line dividing attributes into two categories: visible and invisible (or hidden). The visible attributes are listed above the line (the visible part of the iceberg); the invisible attributes are listed below the line (the hidden part of the iceberg).
A chart showing movement of a value regardless of time, based solely on some time-independent criteria.
Shows the decomposition of some object or system by exposing internal layers sequentially.
A form of visualization of a process or system over time compared to value at each point in time, grouped into four stages: research and development (R&D)/initiation, ascent, maturity, and decline.
A chart that shows ordered points connected by a line to show trends.
A chart with the X-axis showing a list of values within a category (such as a list of business units) where the width is each bar’s relative magnitude compared to the others and the Y-axis showing percentages or ratios. Each value is then a stacked area chart with each area shown as a percentage of the total for that X value. Named for printed fabric patterns typical of Finnish company Marimekko. Also called Mekko chart, matrix chart, eikosogram.
A chart showing movements in a value over time at different points within each time grain, using both lines and bars. This chart shows values for high and low separate from those for start and end of each time period point on the chart.
A form of visualization showing the structure of an organization using trees and levels to show relative hierarchies of teams or individuals.
A form of visualization showing a series of vertical parallel lines representing dimensions or axes, and horizontally oriented lines intersecting points on those exes. During development, the ordering of the vertical axes may need to be shifted to better show patterns in the coordinates.
A chart showing both bars and a line, where the line shows the cumulative total of the individual bars going left to right. Named for Vilfredo Pareto.
A form of visualization using a series of horizontal lines, each representing an evaluation range of a specific quality. A set of processes or performances is evaluated against the lines, and the points of the evaluations are connected into a line for each process or performance.
A form of visualization resembling looking down into a box, with the floor of the box being the main topic, the left and right side panels representing positive and negative input or experiences, the bottom side panel representing prior knowledge or experience, and the top side representing open questions or issues.
A form of visualization for distributed systems, using bars and circles to represent events and conditions respectively. Directional arrows show the path between the events and conditions in the system. Named for Carl Adam Petri. Also called place/transition net, P/T net.
A form of visualization that shows percentages as sectors (slices) of a circle, resembling a pie.
A form of visualization showing a circle with sectors, using radius length of sectors to show relative differences. May have multiple sections to each sector to compare multiple values.
A framework for evaluating strategic positions, using five forces: threat from competitors, threat of substitute products or services, bargaining power of customers, bargaining power of suppliers, and barriers to entry. Named for Michael Porter.
A quartered plot chart used most frequently to determine priorities in business, using growth rate on one axis and market share as the other. Creates four categories: stars (high growth and high market share), cash cows (low growth, high market share), dogs (low growth and low market share), and question marks (high growth and low market share).
A visual representation of how control moves between logical processes (how the end state of a process serves as the start state for other processes).
A visual representation showing three or more quantitative values represented on radial axes of a circle.
A two-dimensional representation of values of a dataset. Usually used to show dependency of one uncontrolled variable vs. another controlled variable.
A form of visualization consisting of vertices (concepts) and directed or undirected edges (relationships).
A representation of the time sequence of objects participating in a process over time. Swimlane charts are a form of sequence chart.
A form of visualization showing trends and variations of multiple measurements over time in one chart.
A form of visualization using a time-varying image that shows the spectral density of a signal over time, using horizontal axis as time, vertical axis as frequency, and hue of the representation as amplitude.
A form of visualization where a project appears in the center and stakeholders are illustrated in terms of proximity of responsibility to the project. Internal stakeholders are above a central line through the project, and external stakeholders are below the line.
A form of visualization using a quartered chart to show stakeholders in terms of importance and influence.
A visual representation of a system where quantities of something travel through the system from point to point over time.
A form of visualization used to document strategic goals from multiple perspectives.
A form of visualization that plots price vertically and quantity horizontally. The supply curve (usually trending upward left to right) shows the price per quantity offered by a supplier. The demand curve (usually trending downward left to right) shows the price per quantity desired by consumers. Equilibrium is the intersection of both curves.
A form of process flow diagram that shows involvement over time within the process for multiple equivalent actors, such as teams, departments, and systems.
A form of visualization that matches goals of different time horizons with the specific technologies necessary to meet or enable those goals.
A form of visualization showing an image with a foundation and two or more pillars supporting a roof, with or without a cloud of distantly related topics. The foundation contains fundamental elements, the pillars group supporting elements, and the roof includes overarching topics that cover all the pillars/groups.
A chart showing a horizontal line or bar containing points labeled with dates and/or events.
A method of representing a hierarchical set of data in a graphical form, with fewer nodes at the either the top (i.e., descendent genealogy) or bottom (i.e., ancestor genealogy).
A chart where one set of values is represented by areas of rectangles and other sets of values are represented by colors. Used to look at large, fast-changing sets of structured data. In this chart, the size of a rectangle reflects its importance, and color conveys urgency (blue shades for positive, red shades for negative). Invented by Ben Shneiderman at the University of Maryland. See also chart, heat map.
A form of visualization showing actors and roles when interacting with objects in defined scenarios. UML and flowcharts are common formats for use-case diagrams.
A form of visualization that shows a problem, the steps to planning a solution, and then the steps to evaluate the results afterward. Shaped like the letter v; hence the name.
A form of visualization that shows all potential logical relationships between a finite set of objects. Used most often to illustrate the concepts of UNION, INTERSECTION, and EXCLUSIVE OR of sets. Named after John Venn.
A chart that shows cumulative effects of sequentially applied values.
A statement of objectives, scope, and stakeholders or participants in a project or program.
A generative pretrained transformer (GPT) artificial intelligence chatbot developed by OpenAI, Inc. This large language model interacts in a dialogue format and can answer questions in natural language. The model imitates human behavior and is trained using reinforcement learning from human feedback.
A synchronization step between a data system and an application where all changes to the data system are recorded to disk and noted as complete.
A copy of the state of a system at a point in time.
A corporate officer who is responsible for managing the enterprise’s data assets.
An executive data steward who serves as the chair of the data governance council and the primary business champion of a data management program.
The head of the information technology (IT) group within an organization. Often reports to the chief executive officer. The prominence of this position has risen greatly as IT has become a more important part of organizations.
An organizational leader responsible for ensuring that the organization maximizes the value it achieves through the organization’s collective knowledge: its intellectual capital (including patents), the skills and experience of its people, the maturity of its processes, and its customer relationships. The CKO is responsible for managing these intangible assets through knowledge management, fostering innovation, sharing best practices, facilitating communication, and avoiding knowledge loss after organizational restructuring.
The executive accountable for discovery and governance of significant risks (strategic, reputational, operational, financial, or compliance-related) to an organization and related opportunities. Data governance is a form of risk management and may be part of this executive’s organization. Also called chief risk management officer.
An executive position focused on technical issues in an enterprise. In technical industries, the CTO heads research and development. In other enterprises, the term is sometimes synonymous with chief information officer (CIO).
A person recognized as a member of a public state, with associated obligations and rights. Not the same as customer.
Solutions for capturing and maintaining accurate, up-to-date data about individual citizens and delivering information in an actionable form “just in time” at citizen touchpoints. A specialized form of master data management focusing on citizen master data.
Establishing relationships with individual citizens and then using that information to treat different citizens differently. Census profiles and taxpayer analysis are examples of decision-support activities that can affect the success of citizen relationships. Effective CRM is dependent on high-quality master data about individuals and organizations; see citizen data integration (CDI).
A measurement that evaluates freedom from obscurity or extraneous data.
A type or category of things with common attributes. Members of a class conform to the definition of the class. Type and category are synonyms for class. Classes are the basis for object-oriented analysis, design, and development, where a class is roughly equivalent to an entity with the addition of described functional behavior. See also method.
A set of objects that share the same attributes, operations, methods, relationships, and semantics.
A class diagram resembles an entity relationship (ER) diagram, but the operations or methods section is not present in an ER. See also chart, class diagram.
A word used in an attribute’s name to show what type of data is contained therein, usually applied at the end. See also prime word.
In .NET framework, developed by Microsoft, associates information with a target element.
In .NET framework, developed by Microsoft, associates information with local system processes.
Represents the security level that can be assigned to users.
The traditional, top-down comprehensive approach to implementing business intelligence, including: building an enterprise data model, defining the data warehouse architecture, designing and constructing the physical database, designing and constructing and testing extract-transform-load programs, and populating the database using current sources and historical data conversions. Used in contrast to incremental data warehouse development.
Generally, a set of discrete, exhaustive, and mutually exclusive observations that can be assigned to one or more variables to be measured in the collation and/or presentation of data.
In data modeling, the arrangement of entities into supertypes and subtypes.
In object-oriented design, the arrangement of objects into classes and the assignment of objects to these categories.
In artificial intelligence, a supervised learning task where the model predicts categorical labels for input data.
Arrangement or division of objects into groups based on characteristics that the objects have in common.
In data management, how data is stored, accessed, and processed by separating responsibilities between two types of systems: clients and servers.
In client/server programming, a software program used to contact and obtain data from a server software program on another computer. Each client program is designed to work with one or more specific kinds of server programs.
A more familiar name for the Information Technology Management Reform Act of 1996, a US federal law coauthored by Congressman William Clinger and Senator William Cohen, designed to improve the way the federal government acquires and manages information technology (IT). It requires departments and programs to use performance-based management principles for acquiring IT and mandates the use of a formal enterprise architecture for all federal agencies.
An architecture in which all access to shared resources is provided on demand via self-service internet applications. Formerly known as distributed computing. The cloud is used heavily in large-scale analytics and in data science applications.
Services that are made available in a distributed computing (cloud) environment.
To store data physically adjacent (in sequence) on a disk, based on a clustering index.
A data science methodology involving finding groups in a given dataset, usually using the distances among the data points as a similarity metric.
Generally, a language-independent set of letters, numbers, or symbols that represent a concept whose meaning is described in a natural language.
In software, the program language lines of instruction that make up software.
In data modeling, a shorthand key value representing the domain value of an attribute. Code sets are intensional domain value sets.
To represent data in a form that can be accepted by a data entry program.
The definition and maintenance of coded data values, descriptions, definitions, cross-references, parent–child rollups, and other relationships for the valid instances of limited (intensional) domains. Code management is a specialized form of master data management. It is a key responsibility of operational data stewards because it has a very significant impact on overall data quality. Code management typically includes an approval process for all code value additions, changes, and retirements.
The definition and maintenance of program code for the purposes of controlling development on production systems.
A relational database table containing rows for each valid value in a finite domain. Code tables contain some form of encoded data values. Code tables are reference data, maintained through code management.
The process of converting verbal or textual information into codes representing classes within a classification system, to facilitate data processing, storage, or dissemination.
The assignment of an incorrect code to a data item.
In software, an error in program lines of instruction that make up software or data transformation routine.
Systems and applications that simulate human processes and mimic the functions of the human brain through self-algorithms using data mining, pattern recognition, and natural language processing.
A close working relationship between parts, complete enough when together to enable some degree of autonomy without other extraneous parts.
A recommendation technique that predicts a user’s interests by collecting preferences from many users.
The assembly of documents or data entities or attributes into a standard order, such as alphabetical.
In data modeling, a data attribute as implemented in a relational database as a vertical component of a table, similar to a field in a flat file record.
In data modeling, supplementary descriptive text that can be attached to data or metadata.
Acronym to identify software that is used as-is, without any customization.
The SQL statement that concludes a unit of work (database transaction).
A single, formal, comprehensive, organization-wide data architecture that provides a common context within which all data is understood, documented, integrated, and managed. It transcends all data at the organization’s disposal and includes primitive and derived data; atomic and combined data; fundamental and specific data; automated and nonautomated (manual) data; current and historical data; data within and without the organization; high-level and low-level data; and disparate and similar data. It includes data in purchased software, custom-built application databases, programs, screens, reports, and documents. It includes all data used by traditional information systems, expert systems, executive information systems, geographic information systems, data warehouses, object-oriented systems, and so on. It includes centralized and decentralized data regardless of where they reside, who uses them, or how they are used.
A standard protocol that defines web server delegation of web-page generation to console applications, known as CGI scripts.
The Object Management Group vendor-independent architecture and infrastructure for object-based programming interoperability. See also Common Object Model.
A form of UML (Unified Modeling Language) diagram that shows the interactions between objects or parts in terms of sequenced messages. Each message is numbered regardless of its placement on the diagram so that the reader can follow the path by following the numbers sequentially.
All professionals and other stakeholders with active interest and role within a community. See also data management community of interest.
The extent to which differences between statistics can be attributed to differences between the true values of the statistical characteristics.
A technology for building and managing information systems. The goal of CEP is to enable the information contained in the events flowing through all of the layers of the enterprise information technology infrastructure to be discovered, understood in terms of its impact on high-level management goals and business processes, and acted upon in real time. This includes events created by new technologies such as radio-frequency identification (RFID).
The act of agreement to follow external government or industry regulations.
The process of conforming, completing, performing, or adapting actions to meet the rules, demands, or wishes of another party.
A content management system that manages low-level objects (image, table, etc.) rather than higher-level documents.
Microsoft Corporation’s programming specification for object interoperability through sets of predefined routines called interfaces.
A model that includes other models and the relationships between them.
The study of using computational methods to process and analyze human language.
The use of software tools (CASE tools) to assist in the development and maintenance of software. All aspects of the software development lifecycle can be supported by software tools, so tools for project management, business and functional analysis, system design, code storage, compiler translation, and testing can all be considered CASE tools. Sometimes more broadly referred to as computer-aided systems engineering.
An old-fashioned term for the management of metadata between the encyclopedias of multiple CASE tools, of the same type or different types.
Automated modeling tools for used model-driven systems planning, analysis, design, and development. “Upper CASE tools” are modeling tools used for planning, analysis, and high-level logical design. “Lower CASE tools” are used for program design, code generation, version control, and testing. Data modeling tools may be both upper and lower CASE tools.
Video, images, and printed media created with computer graphics applications.
A subfield of artificial intelligence that enables machines to interpret and make decisions based on visual data, such as images and videos.
Including only necessary parts; not including unnecessary details or attributes.
The control of process contention for resources within multi-process systems.
The ability of one process to see information about other processes that are executing at the same time.
The space between an upper and lower limit of a range, where there is a high probability of the inclusion of a particular value.
A measurement of certainty that a statistical prediction is accurate.
In artificial intelligence, a probability-like measure indicating how sure the model is about a given prediction.
Ensuring that information is accessible only to those authorized to have access.
In data security, a property of data indicating the extent to which its unauthorized disclosure could be prejudicial or harmful to the interest of the source or other relevant parties.
A generic term often used to describe the whole of the activities concerned with the creation, maintenance, and control of databases and their environments.
Agreement to follow internal policies, standards, procedures, and architecture requirements.
A dimension that means and represents the same thing when linked to different fact tables.
The state of being similar to accepted standards or to the attributes of peers.
The process of becoming similar to the attributes of peers or to a standard.
The characteristic of a graph in which there exists at least one path from every node to every other node in the graph.
The agreement of a group to a decision, judgment, or definition, when all stakeholders present can say, “I can live with it.”
Uniformity or agreement among things or parts of things. Having internal logical and numerical coherence; having no internal contradiction.
The process of combining and aggregating data from different systems and possibly disparate formats to create a unified view of information.
Generally, a restriction on a business action and the resulting data. For example, “Only wholesale customers may place wholesale orders.”
In data management, a specification of what may be contained in data or a dataset in terms of the content or, for data only, in terms of the set of key combinations to which specific attributes (defined by the data structure) may be attached, and how. Examples include dependency (must have at least one), exclusivity (at most one; non-overlapping), subset, or equality.
A type of constraint on an attribute that defines the values that may be assigned, through limits, lists, or ranges.
A type of constraint on a dataset that restricts the combinations of attribute values according to certain rules (e.g., uniqueness).
The information contained within documents and web pages.
The name of a DCMI element set (Coverage, Description, Type, Relation, Source, Subject, Title). See also Dublin Core Metadata Initiative (DCMI).
The processes, techniques, and technologies for the organizing, categorizing, and structuring of information resources so that they can be stored, published, and reused in multiple ways. Content management is a critical data management discipline for data found in text, graphics, images, and video or audio recordings.
A system used to collect, manage, and publish information content, storing it as components or whole documents, while maintaining the links between components. It may also provide for content revision control.
A DAMA International policy stating the organization’s intention to avoid reference to specific technology vendor firms and their products. Also referred to as vendor neutrality.
Generally, facts or circumstances that relate to a situation or event.
In software design, the minimal set of data required for a task that allows interruption and resumption of the task without error.
The process of adding language to signal relevant aspects of an event or data attribute.
A ready state of functionality that seeks to guarantee computing-system operation despite any challenging event. Continuous availability requires seamless availability during any planned or unplanned event and seamless recovery of applications, data, and data transactions committed prior to the event.
Shows the transition of a topic from one extreme to the other, and all interesting points in between. Usually shown on a double-headed arrow, with each end being one extreme.
An element in DCMI element set Intellectual Property: an entity that contributes to a resource. See Dublin Core Metadata Initiative (DCMI).
The mechanism used to maintain acceptable performance of a process.
Data that guides a process, such as indicators, flags, counters, and parameters.
The minimum or earliest acceptable value in a range of acceptable values.
The maximum or latest acceptable value in a range of acceptable values.
A defined list of explicitly allowed terms and their definitions. The organization of a controlled vocabulary into a parent–child hierarchy is a taxonomy.
In systems, the migration from use of one application to another.
In data management, the process of preparing, reengineering, cleansing, and transforming data and loading it into a new target data structure. Typically, the term is used to describe a one-time event as part of a new database implementation. However, it is sometimes used to describe an ongoing operational procedure.
A type of deep learning model particularly effective for image processing and recognition tasks.
An identifier used by a web application to associate a present website visitor with their previous activity with that company.
A style of application processing in which presentation, business logic, and data management are split among two or more software services operating on one or more computers. In cooperative processing, individual software programs (services) perform specific functions that are invoked by means of parameterized messages exchanged between them.
A business data steward with additional responsibility for
- leading data stewardship teams and
- representing data stewardship issues and integrating team models and specifications on a data stewardship committee (DSC).
The set of exclusive privileges granted to an author, creator, or owner of a work and allowing control or use of that work, including copying, distribution, and adaptation of the work.
An architecture promoted by Bill Inmon that describes the complete data lifecycle within an organization through multiple layers and components of architecture in order to satisfy both operational and analytical needs.
A predictive relationship between two factors, such that when one factor changes, the nature, direction and/or amount of change in the other factor can be predicted. Not necessarily a cause-and-effect relationship.
A function that describes the correlation of the values of a dataset to a line.
Comparison of the estimated value of business benefits over time to the estimated cost of expenditures required to realize these benefits.
A DCMI element in element set Content: the topic, jurisdiction, or spatial scope of a resource. See Dublin Core Metadata Initiative (DCMI).
An internet bot that uses search engines to carry out automatic search and retrieval of selected information on behalf of a user. Also called web crawler.
In the context of data, a process of systematically scanning data sources to extract valuable metadata from those sources.
The only functions of data in persistent storage, with a convenient acronym form.
An information value chain analysis tool. It documents that a given organization, role, process, or application creates, reads, updates, and/or deletes data in a given subject area, entity, or attribute. CRUD matrices are the vehicle for information value chain analysis. Each of the many different kinds of CRUD matrices establishes the link between a data model and another model (data-to-process, data-to-organization, or data-to-application).
A DCMI element in element set Intellectual Property: an entity that is responsible for the first existence of a resource instance. See Dublin Core Metadata Initiative (DCMI).
One of the most important prerequisite conditions necessary for an enterprise to reach its goals.
Of interest to more than one organization in an enterprise, especially of data or process.
The practice of suggesting the purchase of a related product to customers who are already making a purchase.
Cross-referencing of data from one or more sources for analysis or reporting.
A medium of exchange, usually a form of money.
Monetary denomination of the object being measured.
The degree to which data represents reality as of a point in time.
A date when the data is considered valid. Also known as currency date or the “as of” date.
A person or organization whose needs are important to the enterprise or person, and whose satisfaction with the products and services provided by the enterprise determines its success, failure, and effectiveness. See also citizen.
Solutions for capturing and maintaining accurate, up-to-date data about individual customers and delivering information in an actionable form “just in time” at customer touchpoints. A specialized form of master data management, focusing on customer master data. See also citizen data integration (CDI).
Establishing relationships with individual customers and then using that information to treat different customers differently. Customer buying profiles and churn analysis are examples of decision-support activities that can affect the success of customer relationships. Effective CRM is dependent on high-quality master data about individuals and organizations; see customer data integration (CDI). See also citizen relationship management (CRM).
Any type of internet-based promotion through websites, targeted email, internet bulletin boards, e-commerce, and online social networking mechanisms.
A metaphoric abstraction for a virtual reality existing inside computers and on computer networks. While cyberspace should not be confused with the real internet, a website might be said to “exist in cyberspace.” According to this interpretation, events taking place on the internet are not therefore happening in the countries where the participants or the servers are physically located, but instead are happening “in cyberspace.”
The time required to execute a process from start to finish.
The former research and education affiliate of DAMA International, with a mission to promote development of a formal, certified, recognized, and respected data management profession.
An international not-for-profit association of data management professionals with chapters and members around the world, dedicated to advancing the concepts and practices of managing data, information, and knowledge as enterprise assets. Founded in 1980 as the Data Administration Management Association, DAMA International is the leading data management professional organization worldwide.
The organizing structure consisting of a functional decomposition of ten data management functions mapped against six environmental factors.
An umbrella term for content that exists on darknets. Darknets use the internet but require specific software, configurations, or authorizations to access. Through the dark web, private computer networks can communicate and conduct business anonymously without divulging their location.
A business intelligence application that consolidates, aggregates, and graphically presents performance measurements compared to goals, arranged so that information can be monitored at a glance. Dashboards can be used to manage any scope of operations.
A representation of facts, concepts, or instructions in a formalized manner suitable for communication, interpretation, or processing by human beings or automatic means.
Individual facts that are out of context and have no meaning by themselves. Often referred to as raw data. The word data is plural in form, but it can be singular or plural in construction.
Subject-oriented, integrated, time-variant, nonvolatile collections of data in support of business intelligence activities. See also online analytical processing.
A dataset created through a computational step applied to atomic data. Derived data is the result either of relating two or more attributes of a single transaction (such as an aggregation) or of relating one or more attributes of a transaction to an external algorithm (formula) or rule. See also data attribute, derived.
Data not structured in a relational database table or grid format. Includes unstructured data, which has different internal structures, but can include links as well as classification tags as part of the tabular data attributes. See also data, unstructured.
Process-oriented, nonintegrated, time-current, volatile collections of data used to support the daily activities of an enterprise. See also online transaction processing (OLTP).
Data that can be described using a discrete domain of vocabulary terms, organized by inherent patterns into semantic groups or entities, presented by context rather than content.
Data that has not been tagged or otherwise structured into rows and columns or records, including documents, files, graphics, images, text, reports, forms, video or sound recordings, social media content, web pages or blogs, engineering and design CAD/CAM files, genomics and scientific sequencing data, and internet of things and machine-generated data. This term has some inaccurate connotations, as there is usually some structure (for instance, paragraphs and chapters) in these formats.
The formal, sometimes highly rigorous, process associated with acknowledging that data has been delivered or accepted for use in an acquiring system or organization.
The degree to which a data attribute value closely and correctly describes its business entity instance (the “real life” entities) as of a point in time.
The collection of processes of identification, selection, and mapping of source data to target data, including detection of source data changes, data extraction techniques, timing of data extracts, data transformation techniques, frequency of database loads, and levels of data summary.
The activity performed to obtain data, or have access to it under either limited or unlimited rights for use.
The organization and management of data in multiple types of storage, including databases, spreadsheets, and image or content management systems.
An individual or organization responsible for specifying, acquiring, and maintaining software for data management, and the security and validation of the contents, including the data dictionary and data models.
The study and presentation of data to create information and knowledge.
A business systems analyst who identifies data requirements, defines data, and develops and maintains data models.
A type of data de-identification where personal identities are removed so individuals cannot be reidentified.
A combination of hardware, software, database management system, and storage that cannot be viewed or inspected and that yields high performance in both speed and storage and makes data access simpler.
Servers built specifically for data transformation and distribution. These servers integrate with existing infrastructure either directly as a plug-in or peripherally as a network connection.
A master data analyst, responsible for the overall data requirements of an organization, its data architecture and data models, and the design of the databases and data integration solutions that support the organization.
The degree to which data models and database designs are stable, flexible, reusable, aligned with enterprise goals, and supportive of data integrity.
The definition and modeling of the information needs of the enterprise and the designs to meet those needs. Includes information needs analysis, enterprise data modeling, definition of related data architecture, and project-related conceptual, logical, and physical data modeling. Physical database design is considered part of database management.
A master set of data models and design approaches identifying the strategic data requirements and the components of data management solutions, usually at an enterprise level. Typically consists of:
- an enterprise data model (contextual/subject area, conceptual or logical);
- state transition diagrams depicting the lifecycle of major entities;
- a robust information value chain analysis identifying data stakeholder roles, organizations, processes, and applications; and
- data integration architecture identifying how data will flow between applications and databases.
The data integration architecture may divide into database architecture, master data management architecture, data warehouse and business intelligence architecture, and metadata architecture. Some enterprises also include
- lists of controlled domain values (code sets), and
- responsibility for the assignment of data stewards to subject areas, entities, and code sets.
The enterprise data architecture is an important part of the larger enterprise architecture that includes business, process, and technology architecture.
The process that supports long-term storage of scientific data and methods used to read or interpret it. Data archival is a step along the path of data preservation and can be phased for online, near online, or offline storage availability. The data archival process is an important part of data migration and data refresh.
A model of delivering data where a provider licenses access via web-based servers for on-demand use.
Any data resource with organizational value.
Most prevalently used in geosciences, the process of combining data samples having specific sample criteria with projected data from a model to create and improve a unified consistent physical system definition.
Data that is written to and contained in static storage.
An inherent fact, property, or characteristic describing an entity or object; the logical representation of a physical field or relational table column. A given attribute has the same format, interpretation, and domain for all occurrences of an entity. Attributes may contain adjective values (red, round, active, etc.).
A unit of data for which the definition, identification, representation, and permissible values are specified by means of a set of characteristics.
A representation of a data characteristic variation in the logical data model or physical data model. A data attribute may or may not be atomic. See also attribute.
The set of possible values for an attribute. The values must conform to the definition of the attribute (such as type or size), and may be expressed by enumeration or by any combination of ranges and individual values, including values and ranges that are excluded from the set.
An instance of an attribute type or domain.
A composite attribute is one that is composed from the concatenation of other attributes.
An attribute created via calculation from some other attribute(s), either within the same object or within a linked or referenced object. See also data, derived.
An attribute or data item that can have multiple values (instances) for an instance of the entity of which it is an attribute. Such an arrangement forms a many-to-many relationship between the entity type and the attribute type, unless each unique value of the attribute can only be associated with at most one instance of the entity, in which case it forms a hierarchical relationship (one-to-many).
An attribute where history is not preserved. All changes overwrite the attribute at the time of the change.
An attribute where all history is preserved by requiring new rows be written that include the new data. The row with the old data is untouched except to update an expiration date or current row indicator.
An attribute where some history is preserved within the same record or row, in separate columns. When new data arrives, old data is moved to other columns within the same row, and the new data overwrites the old data in the column assigned to hold the current value.
The evaluation of data based on defined criteria. Typically used to ensure compliance with contractual and methodological requirements.
The process of artificially increasing the size of a dataset by creating modified versions of existing data.
The extent to which data is accessible when required.
The service-level agreement specifying expected uptime and accessibility for data systems.
The appropriate trade-off between accessibility, quality, security, and the cost of data.
Systematic distortion or imbalance in data that affects fairness or accuracy
An incident that exposes sensitive, personal, or confidential data to unauthorized individuals.
The process by which collected data is put into a machine-readable form.
In relationships, the characteristic of a relationship that specifies the upper and lower bounds of how many instances of one entity or object type can be related to each instance of the same or some other entity or object type. Cardinality is separately specified at each end of the relationship. At each end the choices are 0, 1, or M. Combining the cardinality at both ends of a binary relationship yields 3 × 9 – 1 = 8 possibilities (0:0 is not a valid option).
A curated inventory of data assets with metadata, definitions, and usage guidance.
The process of creating, maintaining, and curating a data catalog.
An organization that values data as an asset and manages data through all phases of its lifecycle.
The process of verifying and stating that a dataset’s contents meet expected standards. See also certification.
Processes ensuring that changes to data or related systems are property documented, tested, and approved.
The attributes of data such as data type, associated metadata, and key data elements.
Activity through which the correctness conditions of the data are verified.
Categorizing data based on sensitivity, importance, or usage policies.
A label indicating sensitivity or security requirements (e.g., Public, Internal, Confidential).
The process of correcting data errors to bring the level of data quality to an acceptable level for information user needs.
A copy of data used for testing, development, or backup purposes.
The process of partitioning the data attributes of an entity or table into subsets or clusters of similar attributes, based on subject matter or characteristic (domain).
Operations performed on data to derive information according to a given set of rules.
The degree to which data is captured.
Compares the attributes implemented in a database against all known requirements.
Algorithms or techniques that change data to a smaller physical size that contains the same information.
The process of changing data to be stored in a smaller physical or logical space.
Protection of data from unauthorized access or disclosure.
The degree to which one set of attribute values matches another attribute set within the same row or record (record-level consistency), within another attribute set in a different record (cross-record consistency), or within the same record at different points in time (temporal consistency).
A person or group that receives data (on a screen, in a report, or through a query) and uses the data to create information. See also information consumer.
The process of changing data structure, format, or contents to comply with some rule or measurement requirement.
The process of changing data contents stored in one system so that it can be stored in another system or used by an application.
Errors, inconsistencies, or damage rendering data inaccurate or unusable.
A person who enters or updates data. Roughly equivalent to data producer. See also create-read-update-delete (CRUD).
A multidimensional data structure that contains an aggregate value at each point (i.e., the result of applying an aggregate function to an underlying relation). Data cubes are used to implement online analytical processing (OLAP). See also schema, star.
The active management of data through its lifecycle to ensure quality and usability.
Often called dedup for short, a feature that can help reduce the impact of redundant data on storage costs. Duplicated portions of the volume’s dataset are stored once and are (optionally) compressed for additional savings.
Statements that specify the business meaning associated with a conceptual, logical, or physical data entity or attribute.
The process of creating business metadata, including names, meanings, integrity rules, and domain values.
In computer programming, the statements in a computer program that specify the physical attributes of the data to be processed, such as location and quantity of data.
The degree to which data definitions are complete, accurate, current, correct, meaningful, thorough, and useful.
Making data broadly accessible to nontechnical users while maintaining governance and controls.
The process of introducing some redundancy into previously normalized databases with an aim of optimizing database query performance.
A relationship in which one data element or process relies on another.
The statements in a computer program that specify the physical attributes of the data to be processed, such as location and quantity of data.
A data model, an architecture model, or a descriptive representation of any complex object.
Analysis, design, implementation, testing, deployment, and maintenance of data.
Any place where business and/or technical terms and definitions are stored. Typically, data dictionaries are designed to store a limited set of available metadata, concentrating on the names and definitions relating to the physical data and related objects of systems implemented or in development. See also repository.
A data dictionary that interacts with its software environment to capture and update metadata in real time.
A store for metadata for multiple software tools. See also repository.
A data dictionary that requires batch or user entry and update of metadata.
The process of locating, identifying, and understanding available data assets.
The secure destruction or deletion of data at the end of its lifecycle.
In data storage, the mathematical patterns of data values as they exist within a set.
In data networks, the patterns of storage of data within and through various systems and on various platforms or sites.
In data movement, transmission of data to one or more locations from a central point.
The use of data mining to uncover relationships in data that may be valid within a test set but are not valid within the wider population. Sometimes used to deliberately generate misleading conclusions. Also called data fishing, data snooping.
Gradual change in data meaning, structure, distribution, or quality over time.
The presence of redundant copies of data.
Activity aimed at detecting and correcting errors, logical inconsistencies, and suspicious data. Data editing is the physical application of data integrity rules, which are developed logically and denormalized within the data to produce data edits, which are then applied to the data.
The process of converting plain text or data into an unreadable format, known as ciphertext, using an algorithm and an encryption key. This ensures only authorized parties with the corresponding decryption key can access and read the original data.
A classification of objects found in the real world – that is, the noun part of speech: persons, places, things, concepts, and events – of interest to the enterprise. Usually expressed in singular form.
An entity or table that resolves a many-to-many relationship between two other related entities or tables.
In a relational model, a child entity of another parent entity that cannot exist on its own.
A diagram that shows the arrangement and relationships between data entities. It contains only data entities and the data relations between those data entities. It does not contain any of the data attributes in those data entities, nor does it contain any roles played by the data attributes.
In software as a service, the practice of keeping a set of data with an independent third party to prevent data loss.
The process of examining data in order to determine ranges and patterns within the data.
A snapshot copy of data from a source database used to update data in a target database, or for use in an application.
To copy data from a source for data movement and data transformation.
The latency of data extracts, such as daily versus weekly, monthly, or quarterly. The frequency that data extracts are needed in the data warehouse is determined by the shortest frequency requested through an order, or by the frequency required to maintain consistency of the other associated data types in the source data warehouse.
The standard expectations of a particular source data warehouse for data extracts from the operational database system of record. A system of record uses an extract specification to retrieve a snapshot of shared data and formats the data in the way specified for updating the data in the source data warehouse. An extract specification also contains extract frequency rules for use by the data access environment.
Software that reads one or more sources of data and creates a new image of the data. See also extract-transform-load (ETL).
An architectural approach integrating data management capabilities across platforms.
A method of transparently joining or linking data from multiple physical locations and/or multiple platforms.
The transfer of data between systems, applications, or datasets.
A visual representation of how data moves or is moved between logical processes or application services (i.e., how the output data from a process serves as the input data for other processes). Essentially a process model, complementary to a data model. Defines the requirements and master blueprint for storage and processing across databases, applications, platforms, and networks.
The exercise of authority, control, and shared decision-making (planning, monitoring, and enforcement) over the management of data assets. See governance; data stewardship.
The highest-tier data governance organization in an enterprise. The DGC includes senior managers serving as executive data stewards, along with the data management leader and the CIO. A business executive may formally chair the council as chief data steward with the data management leader serving as facilitator for council meetings and other activities.
A staff organization of full-time data analysts found in larger enterprises whose mission is to support the data governance council, data stewardship coordinating committees, and data stewardship teams.
Aligning data definitions, formats, and values across systems.
The process of restricting access to data based on concerns regarding proprietary content, economic impact, security implications.
The data that have been identified thus far for potential inclusion in the information system. The process of specifying which data should or will be sought to fulfill user needs. A description of the different types of data and their applicable tools for analysis is also included.
Data that is stored in a distributed network of systems, where the location of the data is unknown and transparent to the user
Data that is carried across networks between systems.
The ability to change the logical or physical structure of data without changing the application program and its view of the data.
On a large scale, the independence of the data architecture from the business activity architecture, the platform architecture, and the information system architecture. On a smaller scale, the independence of the logical design from the physical platform where data will be stored.
1. A single piece of digital information.
2. A specific set of data values for the characteristics in a data occurrence that is valid at a point in time.
3. One data instance is the current instance and the others are historical or for a period of time. Many data instances can exist for each data occurrence, particularly when historical data are instances.
The movement and consolidation of data within and between data stores, applications, and organizations.
An IT professional responsible for data integration processes, practices and software programs across an enterprise.
Part of a master plan for how data is selected, transformed and flows across databases; an important part of enterprise data architecture. It may include database architecture, master data management architecture, business intelligence architecture, and metadata architecture.
A software developer responsible for data integration programming.
Data that complies with all rules regarding definitions, relationships, lineage, and heritage.
In data movement, data that is provably not changed unexpectedly through transmission between systems.
The ability for multiple systems to communicate with one another.
A large storage repository for structured and unstructured data.
Data management architecture that combines data lakes with data warehouses.
The time delay for data to be updated in a system compared to the real world. When data is displayed in real time, data latency is eliminated.
A conceptualization of how data is created and used which attempts to define a “birth-to-death” value chain for data, including acquisition, processing, analysis, storage and maintenance, archiving and retention, and disposal or destruction.
Ability to read, understand and communicate data insights.
The process of populating more than one row at time into database, typically a data warehouse.
The business function that develops and executes plans, policies, practices, and projects that acquire, control, protect, deliver, and enhance the value of data.
A program for implementation and performance of the data management function.
The field of disciplines required to perform the data management function.
The profession of individuals who perform data management disciplines.
In some cases, a synonym for a data management services organization that performs data management activities.
All the data management professionals, data stewards, and other stakeholders with an active interest and role in data management.
One of the business processes within data management.
A generic term used for the highest-level manager of data management services organizations; the manager most directly responsible for data management, including coordinating data governance and data stewardship activity, overseeing data management projects, and supervising data management professionals.
A centralized data management organization performing data management functions; one or more units of data management professionals responsible for data management within an organization. A centralized data management services organization is sometimes known as a data management center of excellence.
Selected courses of actions setting the direction for data management within the enterprise, including vision, mission, goals, principles, policies, and projects.
The assignment of source data entities and attributes to target data entities and attributes, and the resolution of disparate data.
A term used for classifying data at a deep meaningful level for its sensitivity (secret, etc.) and appropriate release. For example, some data will not be sensitive on its own, but will not be releasable to certain countries, or in combination with other data, which then makes it sensitive. Classification and security considerations often shape how markings are applied.
A decision-support database supporting business intelligence in a limited subject area, using a dimensional data model design. Typically, data marts source their data from an enterprise data warehouse or operational data store.
A method for creating a version of data by obscuring individual items within a database. Used to make the data anonymous. Typically used for development or testing.
A decentralized data architecture that organizes data by a specific business domain. It helps scale analytics adoption beyond a single platform and a single implementation team.
An architecture that enforces data security policies within and between domains to provide centralized monitoring of data sharing functions. It enables domain teams to perform cross domain analysis with operational and analytical data models as data products.
A central data management framework that governs and records the data available in an organization. A data mesh model prevents data silos from forming within specific business domain systems.
The process of transferring data from one database to another. See also conversion.
The process of sifting through large amounts of data using pattern recognition, fuzzy logic, and other knowledge discovery statistical techniques to identify previously unknown, unsuspected, and potentially meaningful data content relationships and trends. See also predictive analysis.
A model that includes formal data names, comprehensive data definitions, proper data structures, and precise data integrity rules. A complete data model must include all four of these components.
The visual presentation of the structural portion of a data model with icons for entity records and lines between them to represent relationships. See also data structure diagram.
A data model that is presented at a high level of abstraction, hiding the underlying details, and making it easier for people to comprehend. A conceptual model should reflect the way the users think. For example, many-to-many relationships are common in conceptual models.
A data model that represents data in a star-like structure of only one-to-many relationships, where each entity has either all relationships having the “one” side or the “many” side. See also schema, star.
A conceptual data model or logical data model providing a common consistent view of shared data across the enterprise, however that is defined, at a point in time. It is common to use the term to mean a high-level, simplified data model, but that is a question of abstraction for presentation. See also data model, conceptual (CDM).
A data model that represents data in a tree-like structure of only one-to-many relationships, where each entity may have a “many” side when related to a parent, and a “one” side when related to a child. See also structure, hierarchical.
An entity-relationship data model including data attributes that represent the inherent properties of the data, including names, definitions, structure, and integrity rules, independent of software, hardware, volumetrics, frequency of use, or performance considerations.
A representation of objects and their participation in one or more owner-member sets. In such a model, a both owners and members may participate in multiple sets, affecting a network of objects and relationships.
The definition or representation of a data model for implementation and realization in a particular database management system, including naming convention and physical data types. It may be denormalized for performance and access simplification. A high-level description of a database design without specific physical layout (how the data is stored on disk) information.
A conceptual data model that provides structure and defines meaning for non-tabular data, making that meaning explicit enough that a human or software agent can reason about it. See also ontology.
An analysis and design method, building data models to
- define and analyze data requirements,
- design logical and physical data structures that support these requirements, and
- define business metadata and technical metadata.
The act of creating a data model.
A conventional system of symbols for modeling data.
One of several notation conventions for modeling data, developed by Richard Barker and others in 1986.
The formalism, technique, style, notation, etc. used to guide the data modeling activity to create a data model. The scheme encompasses the notational conventions (syntax) for representing the semantics of the data model. Examples of data modeling schemes include single file, flat file, multi-file, relational, hierarchical, record-based, ER, EER, fact oriented (no file), IE, UML, IDEF1X, Barker, multidimensional, star, and object-oriented.
A data modeling scheme that is not record-based – it does not use the constructs of file or entity tables, and consequently does not suffer from table think. See also fact-oriented modeling; object–role model.
A data modeling scheme in which records are formed by clustering attributes. Used to represent an entity type population. One or more attributes may serve as an identifier. In general, a record could have a hierarchical structure, thus allowing multi-valued data items and repeating groups of items.
A data modeling scheme that uses the constructs of entity, attribute, identifier, and foreign key to represent relationships. Attributes are represented by data items or columns. Entities are represented by tables, which is a cluster of one or more attributes. Relationships between entities are represented by foreign keys. The major distinguishing characteristic of a relational data model is that all attributes must be single valued (atomic). See also flat file.
The process by which businesses earn revenue from their data assets.
The process of extracting data from one system and loading it onto another system. See also extract-transform-load (ETL).
The process that brings data into a normal form that minimizes redundancies and keeps anomalies from entering the data resource. It provides a subject-oriented data resource based on business objects and events.
A deluge of data coming at a recipient that is not relevant and timely. An overabundance of unwanted noninformation.
An individual responsible for definitions, policy, and practice decisions about data within their area of responsibility. For business data, the individual may be called a business owner of the data.
Short statements of management intent and fundamental rules governing the creation, acquisition, integrity, security, quality, and use of data and information.
The process that involves checking or logging the data in; checking the data for accuracy; entering the data into the computer; transforming the data; and developing and documenting a database structure that integrates the various measures. This process includes preparation and assignment of appropriate metadata to describe the product in human-readable code or format.
The degree to which information products (reports, screens, charts) are easy to understand and use without misinterpretation by the intended audience.
The limitation of data access to only those authorized to view the data. See also confidentiality.
The operation performed on data, through capture, transformation, and storage, in order to derive new information according to a given set of rules.
A person, organization, or software service creating or providing data. See also data creator.
A collection of statistics about a data attribute that shows patterns of usage, patterns of contents, and any other patterns that may be of interest.
An approach to data quality analysis, using statistics to show patterns of usage and patterns of contents, and automated as much as possible. Some profiling activities must be done manually, but most can be automated.
The distribution of data from one or more source databases to one or more local target databases, according to defined rules. Typically used in reference to distributed databases. See also data replication.
Information about the origin of an organization’s data.
The degree to which data is accurate, complete, timely, consistent with all requirements and business rules, and relevant for a given use. See also information quality (IQ).
The evaluation of data quality; the identification of inaccurate, incomplete, inconsistent, and untimely data and its causes.
Quality inspection and review processes for data, data models, and database designs, including testing for application data controls and edits, testing of data integration programs, data quality analysis, and data quality audits. See also quality assurance (QA).
The random sampling of data and testing it against its valid data values to determine its accuracy and reliability.
A declaration that a database has met a set of defined data quality requirements (service levels), based on data quality analysis and audits.
The rate that a data attribute loses accuracy over time if not updated. For example, if age of individuals is stored, the entire dataset’s age attribute will be incorrect after one year because everyone will have aged by then.
The application of total quality management concepts and practices to improve data and information quality, including setting data quality policies and guidelines, data quality measurement (including data quality auditing and certification), data quality analysis, data cleansing and correction, data quality process improvement, and data quality education.
System analysis and redesign to eliminate or prevent data errors and defects. A proactive, preventive approach to improving data quality.
Application requirements that eliminate or prevent data errors, including requirements for domain control, referential integrity constraints, and edit and validation routines.
The process of adjusting data derived from two different sources to remove, or at least reduce, the impact of differences identified.
A physical grouping of data items that are stored in or retrieved from a data file. It is referred to as a row or tuple in a relational database. A data record represents a data instance.
The storing of the same data in multiple places, either intentionally (for backup) or unintentionally (leading to inconsistent data).
The unknown and unmanaged duplication of business facts.
The process of analyzing, standardizing, and transforming data from nonstandard files and databases into a standardized database that is part of the enterprise data architecture.
The process of applying updates as a group to a dataset, then allowing users access to the updated data.
The residue of data that has been nominally erased or removed.
The consistent copying of data from one primary data site to one or more secondary data sites. The copied data is kept in synch with the primary data on a regular basis.
Statements describing the data needs of a person or organization. Business metadata (data names and meanings) and logical data models are structured ways of defining data requirements, in addition to more traditional requirement specifications.
Potential negative impact from data misuse or failure.
The safety of data from unauthorized and inappropriate access or change.
The measures taken to prevent unauthorized access, use, modification, or destruction of data.
The process of ensuring that data is safe from unauthorized and inappropriate access or change. Includes focus on data privacy, confidentiality, access, functional capabilities, and use.
An interface to a business process that receives or delivers data attributes, usually via a web application.
Exchange of data and/or metadata in a situation involving the use of open, freely available data formats, in which process patterns are known and standard, and the exchange is not limited by privacy and confidentiality regulations.
An agreement between parties that describes the allowed activities, uses, and restrictions regarding data shared between the parties.
A collection or repository of data that is controlled by one department or business unit and not easily accessible by the rest of the organization.
Origin from which data is collected.
Connection information for a database used in an Open Database Connectivity connection setup.
The process of moving data from one system into intermediate storage before final processing into a target.
A database that stands between the operational source databases and the target databases (typically an operational data store, data warehouse, or data mart). The data staging area is considered the “back room” portion of the data warehouse environment. It is where the extract, transform, and load effort takes place, and is out of bounds for end users. Most data in the data staging area is transient, although typically there is some relatively small amount of persistent data.
The agreed-upon format or definition for data elements.
A business leader and/or subject matter expert designated as accountable for:
- the identification of operational and business intelligence data requirements within an assigned subject area;
- the quality of data names, business definitions, data integrity rules, and domain values within an assigned subject area;
- compliance with regulatory requirements and conformance to internal data policies and data standards;
- application of appropriate security controls;
- analyzing and improving data quality; and
- identifying and resolving data-related issues.
Data stewards are often categorized as executive data stewards, business data stewards or coordinating data stewards. See also data owner; data stewardship; data governance.
Formal, specifically assigned, and entrusted accountability for business (nontechnical) responsibilities ensuring effective control and use of data and information resources. See also data steward; stewardship; data governance.
Formal accountability for business responsibilities ensuring effective control and use of data assets.
A permanent cross-functional group of coordinating data stewards responsible for
- supporting the data governance council, and
- integrating the work of data stewardship teams.
The data governance council may delegate responsibilities to the DSC. The DSC may be led by the chief data steward, data management leader, and/or the enterprise data architect. A large organization may have additional DSCs at one or more levels lower than the enterprise DSC.
A generic term for any committee of data stewards. May be synonymous with data governance council (DGC), with both executive data stewards and coordinating data stewards, or may be a working committee of coordinating data stewards below the DGC.
A temporary or permanent focused group of business data stewards collaborating on data modeling, specification, and data quality improvement, typically in an assigned subject area, led by a coordinating data steward and facilitated by a data architect.
The means of recording or archiving data so that it is available for future use.
A place where data is stored; data at rest. A generic term that includes databases, flat files, and nonelectronic data files.
A business plan for leveraging an enterprise’s data assets to maximum advantage. See also enterprise data strategy.
A set of structural metadata associated to a dataset, which includes information about how concepts are associated with the measures, dimensions, and attributes of a data structure along with information about the representation of data and related descriptive metadata.
The visual presentation of a data model with icons for entities with attributes and lines between them to represent relationships and roles. First known published article to propose such diagrams was by Charles Bachman in 1969. See also data model diagram; entity-relationship diagram (ERD).
A person, place, thing, concept, or event that is of interest to the organization and about which data is captured and maintained in the organization’s data resource.
An unmanaged data lake that is either inaccessible to intended users or provides little value. Data swamps occur when there is little data governance and/or proper planning.
The continuous harmonization of data attribute values between two or more different systems, with the end result being that data attribute values are the same in all of the systems.
The process of combining and evaluating data from different sources. The aim is to identify patterns and relationships in the data for analysis.
The evaluation, selection, implementation, inventory, and maintenance of hardware and software products in support of data management, including database management systems, data modeling and database administration tools, metadata management tools, data integration tools, and business intelligence tools.
The process of moving data from one system or operating environment to another.
Changing the format, structure, integrity, and/or definitions of data from the source database to comply with the requirements of a target database.
A category of physical data structures with common physical properties and uses, such as numeric, alphanumeric, packed decimal, floating point, and datetime.
A set of distinct values characterized by properties of those values and by operations on those values.
The process of monitoring the results of data compilation and ensuring the quality of the computational results.
The process of evaluating stored data against established acceptance criteria to determine its quality and usability.
The specific representation of a value for an attribute as of a point in time.
The flow of data across processes in support of the enterprise’s business value chain.
The identification of which functions, processes, applications, organizations, and roles create, read, update, and delete (CRUD) different kinds of data (subject areas, entities, attributes), expressed in CRUD matrices, particularly when the compared items are arranged in value chain sequence.
A concept, originated by Dan Linstedt, that organizes data for decision support into hubs, links, and satellites. This method supports rapid and flexible loading of warehouse data.
Accessing data without physical movement.
Techniques for graphical representation of trends, patterns, and other information.
The amount of data stored or processed.
An integrated, centralized decision-support database and the related software programs used to collect, cleanse, transform, and store data from a variety of operational sources to support business intelligence. A data warehouse may also include dependent data marts.
A subject-oriented, integrated, time-variant, and nonvolatile collection of summary and detailed historical data used to support the strategic decision-making processes for the corporation.
A copy of transaction data specifically structured for query and analysis.
A system containing integrated servers, storage, and software specifically optimized for data warehouse processing.
A collection of checks and balances to ensure the right data is extracted from the right sources, then transformed, cleansed, and summarized correctly, and finally loaded to the right target database tables.
A relational database management system or multidimensional database management system. Data warehouse engines require strong query capabilities, fast load mechanisms, and large storage requirements.
An integrated network of data warehouses that contain sharable data propagated from a source data warehouse based on information consumer demand. The warehouses are managed to control data redundancy and to promote effective use of the sharable data.
Implementation of a data warehouse that enables real-time or near-time data analysis.
A conceptual data warehouse made up of multiple decision-support databases, potentially on multiple servers, but presented transparently to business intelligence users as a unified schema for query, analysis, and reporting.
An enterprise data warehouse fed by extracts from departmental data warehouses and/or legacy data warehouses prior to their incorporation and/or retirement.
A data warehouse that draws data from nearby operational systems and supports a distinct organization, functional area (such as manufacturing), or geographic unit within the enterprise. Distinct from an enterprise data warehouse.
The operational extract, cleansing, transformation, and load processes, and associated control processes, that maintain the data containing within a data warehouse.
The storage of data for the analysis of trends and patterns in the business.
The operational, administrative, and control processes that provide access to business intelligence data and support to knowledge workers engaged in reporting, query, and analysis.
A US government website launched by the federal chief information officer of the United States in order to make government-collected data available for use by the public.
An organized collection of data stored in a structured way to enable rapid search and retrieval by a computer.
A common application programming interface used to call database functions to allow an application to connect to multiple different databases without function calls for all possible databases.
A database structure that serializes values by columns and then by rows, rather than conventional databases, which serialize values by rows and then by columns.
A database management system that stores content by column rather than by row.
A database management system that is data model independent and designed to efficiently handle unplanned, ad hoc queries in an analytical system environment. Unlike relational database management systems (records-based storage) or column-oriented databases (column-based storage), a correlation database uses a value-based storage architecture in which each unique data value is stored only once and an auto-generated indexing system maintains the context for all values.
A database that contains objects residing on independent systems in a network, but can be accessed as though all objects reside on the same system.
A database in which all relationships among data entities and attributes are hierarchical (see relationship, hierarchical). Sometimes used to refer to databases that have hierarchical record structures but allow more general network relationships between record types. Examples include UML, most object-oriented data models, Kroenke’s Semantic Object Model, the SQL:99 standard (and later versions), and nested relations. See also structure, hierarchical.
A data structure with three or more independent dimensions.
A type of database in which records are stored with links or pointers to other records. Distinguished from a hierarchical database in the sense that a child record may have relationships with multiple parent records.
A database based on the object-oriented paradigm instead of the relational model.
A database supporting one or more transactional applications. Operational databases are the sources of data for operational data stores and data warehouses. They contain detailed data used to run the day-to-day operations of the business. The data continually changes as updates are made. See also online transaction processing (OLTP).
The most common form of database, storing data in tables made of up of columns and rows, with primary keys, created using the relational data modeling scheme.
A database that feeds into a target database. May be an operational database, operational data store, data staging area, or data warehouse.
A database with built-in time aspects, including valid time and transaction time.
The function of managing the physical aspects of data resources, including database design and integrity, backup and recovery, performance and tuning, generally within the context of a particular database management system.
A database administrator (DBA) who supports application systems, sometimes focusing on development and test environments, database design, and SQL tuning as opposed to the entire application stack. Contrasts with operational DBA and procedural DBA.
A database administrator (DBA) focused on support of production environments, including performance tuning, backup and recovery, high availability, job scheduling, data delivery, and security access levels. The operational DBA has a specific responsibility for change management regarding implementation.
A database administrator (DBA) who specializes in development and support of procedural logic controlled and executed by the database management system: stored procedures, triggers, and user-defined functions.
A database administrator (DBA) who supports the development of data management applications.
The process of developing a physical data model, followed by definition of all physical database objects, including tables, indexes, and sequences.
The physical data model and the detailed data definition language for a database. The database design addresses physical constraints such as storage and performance.
The degree to which data in a database conforms to logical integrity constraints through the implementation of physical database management system constraints.
The degree to which data in a database can be recovered in the event of a hardware or software failure.
A comprehensive list of all databases within a system or an organization.
The development and support of structured data resources. Database management is broader in scope than database administration, including responsibilities beyond those of database administrators.
The computer software program used to manage and query a database.
A type of database management system in which parent–child relationships are established between data segments.
A specialized database management system that supports online analytical processing, enabling users to analyze large amounts of data. An MDDBMS captures and presents arrays of data that can be arranged in multiple dimensions.
Database software that stores data as objects, instead of storing the data in a relational format and then instantiating objects in memory.
An object-relational database management system to manage an object-relational data structure, which is a hybrid of object-oriented and relational.
Database management software controlling the creation, storage, manipulation, access, and performance of relational databases.
The use of information about customers and prospects to strengthen customer relationships by identifying new opportunities and improving customer service. Uses methods for creating, testing, and executing marketing strategies based on analysis of customer data. Includes the mass customization of marketing campaigns to decrease costs, improve response, build customer loyalty, reduce attrition, and increase customer satisfaction.
In a distributed application architecture, the database management system software, related data integration and access services, and associated hardware supporting access to and manipulation of data, separate from application logic and user interfaces.
A unit of work; a set of statements to read, create, modify, or delete business data. The database management system must complete performance of all the statements or reverse the changes.
Any organized collection of data.
A point in time with the granularity of a day.
A class word, abbreviated usually to dt.
A DCMI element in element set Instantiation: a timeframe for an event. See also Dublin Core Metadata Initiative (DCMI).
A scenario where a set of multiple simultaneous actions within a set waits for others within the set to complete and release the resources being held. The waiting processes are “locked out” from the resources held by the other processes. A true deadlock lasts forever; it is never resolved.
A subset of SQL used to define data security, user function permissions, and data access to data in relational tables.
Generally, the subset of SQL commands used to define and implement structured database objects.
In database management systems, the specific definitions to formally define and implement a database.
A middleware protocol and application programming interface standard for data-centric communication, designed to facilitate the exchange of data between different systems, applications, or devices in real time, ensuring efficient, scalable, and reliable data sharing in distributed environments.
Recorded information used to diagnose system issues.
In data governance, information about the who, when, and how of making a data-related decision.
An application that uses data to support managerial decisions through ad hoc query, summarization, drill-down analysis, trend analysis, exception identification, and “what if” scenario modeling. See also business intelligence (BI).
Describes a type of programming language in which the programmer does not define the flow of control at execution time.
The process of reversing encryption; decoding back into original format.
The process of reasoning from one state to another, such as from cause to effect, or from general to specific.
Subtraction.
The process of elimination of redundant copies of data from storage or during a merge of multiple datasets.
A specialized form of machine learning that uses neural networks with many layers (deep neural networks) to analyze various types of data, such as images, audio, and text.
A data value that is automatically assigned when no other data values are selected or applied. Also called default value.
A data value that does not conform to its quality requirements. See also error.
The percentage of data that is incorrect, inaccurate, or no longer true. The number of defects found compared to the total number of data values.
The US federal funding source for development of the Semantic Web.
A Six Sigma process improvement method, used for projects to design and create new product or processes.
A Six Sigma process improvement method, used for projects to improve existing business processes.
A statement conveying a fundamental character or the meaning of a word, phrase, or term. A clear, distinct, detailed statement of the precise meaning or significance of something.
The process of assigning names, description, and specification.
To remove or erase.
A SQL statement (command) that specifies removal of data in a relational database.
A Greek letter (Δ) signifying the difference between two statistical values.
The term used to identify rows that have changed between time periods, used in extract-transform-load processing.
A dataset containing only the data that was updated between the last extraction or snapshot process and the current execution of the extraction or snapshot.
The “plan-do-check-act” cycle of continuous improvement developed by Walter Shewhart and popularized by W. Edwards Deming. See also Shewhart cycle.
A segment of a population delineated by certain shared inherent characteristics.
The study of human populations through statistics.
Restructuring data from a nonredundant (normalized) form to a redundant form that is optimized for the decision-support system.
In a relationship, a constraint between any two attributes where one attribute value matches to one and only one value of the other attribute.
Used in the context of an attribute in an entity record, an attribute instance cannot exist without being related to an entity instance (the dependency part) and there can be at most one instance (value) for that attribute for each entity instance (the function part where A = fn(X) : given a value for X, fn uniquely determines a value for A. In this case, X is called the determinant of A).
In a relationship, a constraint between sets of attributes where the values of one set of attributes match to one and only one other set of attribute values. Contrast with functional dependency, where the constraints involve only one attribute from each set.
A type of dependency in which the value of a non-key field is determined by a part of a composite key, thus violating second normal form.
A type of dependency in which the values of non-key attributes are determined by another non-key attribute, rather than the entity key. It is a functional dependency, but between two non-key attributes, hence violating third normal form. See also dependency, functional.
A data mart whose tables are sourced from an enterprise data warehouse.
The act of putting information technology into productive use. Installation puts the system into the production environment. Deployment includes installation, but also includes efforts to train and encourage effective use.
A textual representation of a thing.
A class word, abbreviated usually to desc.
A DCMI element in element set Content: a textual, tabular, or graphical portrayal of a resource. See also Dublin Core Metadata Initiative (DCMI).
Analysis summarizing historical data.
A model that describes how a system actually works.
A deliberate, purposeful plan, layout, delineation, arrangement, and specification of the component parts and interfaces of a product or system. A logical design is an abstract design for fulfilling requirements without consideration for physical constraints. A physical design considers the requirements along with physical constraints.
To conceive, plan, define, arrange, and specify a product or system.
A process where all aspects of a system design are reviewed publicly before code construction starts.
The entity domain that determines the value of an attribute. See also dependency, functional.
A type of matching that relies on defined patterns and rules for assigning weights and scores for determining similarity.
A person who designs, codes and/or tests software. Different types are known as software developer, systems developer, application developer, software engineer, or application engineer.
The measure of difference between expected and observed values, or more generally, between any two values.
A visual representation of relationships between multiple things (i.e., how a system works, how parts are related to the whole). See also chart.
A subset of language used or agreed to by a group of people. An ontology defines the precise meaning of the vocabulary in a dialect and the relationship between these terms.
A collection of definitions for words, terms, and phrases that differentiate closely related words. See also data dictionary.
A digital representation of a topography or terrain, using regular shapes (squares or triangles) to approximate roughness of a surface.
The management of data on digital media over time. As digital media storage mediums and storage applications change, either data must be moved to new media or old storage retrieval mechanisms and applications must be kept operational.
A generic term for technologies that can be used to impose limitations on usage of digital content and/or devices.
To convert something into a binary representation for computer storage and/or use.
Generally, an axis from which one can regard or summarize something.
In architecture, one of a series of properties that together are used to uniquely identify a location or a component of a system.
In business intelligence, a category for summarizing or viewing data (e.g., time period, product, product line, geographic area, organization).
In dimensional modeling, a type of table, or a structural attribute of a data cube containing a list of members, all of which are of a similar type in the user’s perception of the data. For example, all months, quarters, years, etc., make up a time dimension; likewise, all cities, regions, countries, etc., make up a geography dimension. A dimension acts as an index for identifying values within a multidimensional array. Dimensions offer a very concise, intuitive way of organizing and selecting data for retrieval, exploration, and analysis.
In dimensional modeling, a table containing a row for each occurrence of a dimension list, linked to one or more fact tables through use of the dimension table key as a foreign key in each related fact table.
A dimension that exists once but is used in multiple star schemas, so that the dimension content and meaning are the same regardless of which fact table is joined.
A dimension where there are no valid dimensional attributes other than a unique identifier in a one-to-one relationship with a fact table.
A dimension that consists of multiple loosely related codes and indicators collected into one table in order to reduce the number of keys and indexes needed in a star schema.
A dimension that includes attributes of another dimension that change over time more frequently than is desired. This sometimes greatly enhances load performance by concentrating write operations to a small subset of attributes.
A dimension that slowly increases in the number of rows.
A dimension in which all attributes are type 1, so all attributes are overwritten with new data.
A dimension in which all attributes are type 2, so any attribute that changes for the business key requires generation of a new row.
Similar to a type 2 dimension, a type 2A writes a new row for any change in the data for the row and is time-date stamped. However, the old row is retired to a history table; it is not left in the current table. The type 2A table resembles a type 1 in its contents. However, it can be joined to the history table to get a full type 2 view.
A dimensionin which all attributes are type 3, so that any attribute that changes for the business key will require copying attribute values to other attributes before the original attributes are overwritten.
A dimension table where the data is physically split into two tables: one with the current value rows, and the other with only historical value rows.
A dimension table that combines the attributes of types 1, 2 and 3 (1+2+3=6).
A computed value derived from the calculation of a fact measure at the intersection of one or more dimensions at non-granular levels.
A specialized type of physical data model particular to a retrieval-only database design, commonly used in data warehouses and data marts, where denormalized fact tables are linked to dimension tables. Star schemas and snowflake schemas are examples of dimensional models.
A protocol for data access across systems using data files instead of blocks.
A type of storage device where access is directly to the device, rather than through a cache or other interface.
A storage device attached to a server or workstation without the use of a storage network.
Generally, information heavily optimized for searching and reading.
In data storage, a table, index, or folder containing addresses and locations of data or relationships between data objects.
In operating systems, including Windows, a synonym for a folder, used to organize stored files and other folders.
A type of metadata store that limits the metadata to the location or source of data in the enterprise.
Data with a high degree of inaccuracy, incompleteness, or inconsistency, or that fails some edit criteria.
To clarify the meaning of a term by selecting between alternate interpretations.
The process of identifying attributes to differentiate or clarify between alternate interpretations.
A protocol and associated execution to recover lost computing-system usage (applications, data, and data transactions) committed up to the moment of system loss.
Fundamentally dissimilar in kind, or containing or including dissimilar or unlike attributes. Opposite of similar.
Data that are distinctly different in kind, quality, or character. They are unequal and cannot be readily integrated to meet the business information demand.
Processing across multiple coordinated machines.
A method of training machine learning models across multiple machines or devices to scale processing power.
An International Business Machines (IBM) architecture for coordinating data across multiple relational database management systems.
Transaction spanning multiple systems requiring atomic completion.
A hierarchical naming system, built on a distributed database, to associate various information with domain names assigned to each of the participating entities. The DNS serves as the phone book for the internet by translating human-friendly computer hostnames into IP addresses. This system is managed by Internet Corporation for Assigned Names and Numbers (ICANN).
Generally, any information delivery vehicle, paper or electronic.
In data management, the content and structure in an electronic file.
In document or record management, a paper object in the real world, which may include signatures.
Managing data found outside of standard structured databases.
The storage, inventory, and control of electronic and paper documents.
An application used to track and store electronic documents and/or images of paper documents. Document management systems commonly provide storage, versioning, metadata, and security, as well as indexing and retrieval capabilities.
A platform- and language-neutral application programming interface that allows programs and scripts to dynamically access nodes in an XML document and update the content, structure, and style of these documents.
Descriptive text and images used to define or describe an object, design, specification, instructions, or procedure.
Generally, a set of things that have a common definition, such as the set of possible values for an attribute or the population of an entity.
In data modeling, a type of attribute with common properties and purposes, such as key, code, date, indicator, amount, name, or description.
In an ontology, a constraint limiting the classes that can use a property.
A characteristic of multiple attributes using a domain where the domain of valid values used are not internally consistent from attribute to attribute, or are not applied consistently. Example: a unit of measure code domain where one attribute uses the code to show quantity on hand as doz, and another shows reorder point quantity in numerals.
The study of a domain of values for a data item, to determine whether that item is similar to another item and a candidate for integration or merging.
An internet-based company that relies on digital technology and the use of the web as the primary communication and interaction media.
A general condition wherein users cannot use or access computing systems, applications, data, or information for a broad variety of reasons.
The ability to drill down to any dimension without having to follow predefined drill paths.
A method of exploring detailed data that was used in creating a summary level of data. Drill-down levels depend on the granularity of data within a dimension.
An OLAP function often used to imply the ability to navigate from dimensionally aggregated data to relational transaction source data. Typically, the transaction set returned is constrained by multiple filters in accordance with the starting dimensional aggregate.
Data analysis performed on a dataset with applied mathematical functions, associated with fewer dimensions, higher levels of hierarchy in one or more dimensions, or both.
A standard core ontology for metadata about documents, originating in Dublin, Ohio, and managed by the Dublin Core Metadata Initiative (DCMI).
Describes a system that has communication paths in both directions between two parties.
An advanced data warehouse architecture that includes the lifecycle of data in the data warehouse, the integration of unstructured data, and enterprise metadata.
A data dictionary that an application program accesses at run time.
Changing the appearance of the data seen by the end user or system without changing the underlying data. Often used when users need to access sensitive production data.
A method used to solve complex problems by breaking them down into simpler sub-problems and solving each sub-problem only once.
Dynamically constructed SQL queries that are not preprocessed, and whose access paths are determined at run time prior to execution.
Doing business electronically, usually over the internet. The two main types of e-business are business-to-consumer (B2C) and business-to-business (B2B).
Doing business with a commercial enterprise electronically and without other human intermediaries. See also business-to-consumer; e-business.
An 8-bit character encoding used by International Business Machines Corporation (IBM). See also ASCII (American Standard Code for Information Interchange).
A value-based metric for performance measurement, value-based planning, and incentive compensation developed by Joel Stern and G. Bennett Stewart III. EVA is calculated by deducting a charge for the cost of capital from operation profits.
In graph theory, a connection between two nodes in a graph. Also known as an arc.
Artificial intelligence (AI) processing that happens locally on hardware devices (“the edge”) rather than relying on cloud-based computation, enabling faster responses and data privacy.
Ensuring data is created in conformance with business rules. Database integrity controls and software routines can enforce business rules.
Standards-driven technology for high-volume B2B e-business transaction exchange, linking application systems across enterprises so that a transaction on one system at one company generates a like transaction on a system at another company. More sophisticated EDI implementations transform the way business procedures are executed to gain optimal productivity.
The international electronic data interchange standard, developed by the United Nations and maintained by the United Nations Center for Trade Facilitation and Electronic Business. EDIFACT provides a set of syntax rules for data structures, an exchange protocol for interaction, and standard communication messages that allow multicountry and multi-industry exchange. Can be compared to XML; however, EDIFACT data is very cryptic, whereas XML is human-readable.
A system used to track and store electronic documents or images of paper documents.
The systematic collection of individual patient or population health information in digital format.
A strategic initiative of the US National Archives and Records Administration (NARA) to preserve and provide long-term access to federal, presidential, and congressional records.
A sentence having one predicate (verb phrase) and one or more nouns or noun phrases that serve as subject(s) or objects; can also include modifiers with the verb or nouns, such as must, only, at most one, or at least one. The basis for object–role modeling or fact-oriented modeling (formerly NIAM).
The process of providing results of one system using another different system, such that the results are identical even if the processes are not.
A method of communication protocol design that separates network functions from underlying structures.
In object-oriented design, the combination of structure (data and values) and operations (processes; program code) associated with an object. The processes use the data to act on objects.
The conversion of a recognizably meaningful character stream to an unrecognizable character stream by means of a cipher code, in order to secure data and prevent unauthorized access to personally identifiable information and confidential information.
An encryption method in which both writer and reader use the same key to respectively encrypt and decrypt messages or character strings.
An encryption method in which the key to encode the information is different from the key to decode the information.
A non-definable metadata store used by an application development tool.
The scope of an organization as defined by that organization, based on a purpose or point of view. An enterprise may be a business, not-for-profit, government agency, or educational institution. An enterprise has a purpose, goals, and objectives.
Technology that allows data sharing between unrelated systems in an organization, providing a single point of interface to which all applications and databases connect, resolving differences between systems, triggering processes, and delivering data in the proper format to the proper destination.
A web-based approach to distributing business information, consolidating business intelligence objects (reports, documents, spreadsheets, data cubes, etc.) and making them easily accessible, subject to security authorization, to nontechnical users via standard browser technology.
Data that is shared across more than one function within an enterprise, or is created and used by one function but still considered essential to the enterprise.
See data architecture, enterprise. See also architecture, data.
A data layer that separates data sources from applications, providing the means to solve the potential gridlock prevalent in distributed environments such as grid computing and service-oriented architecture.
A structured program for managing physical data resources as they are used by the enterprise.
The development of a common consistent view and understanding of data entities and attributes, and their relationships across the enterprise.
A centralized data warehouse designed to service the business intelligence needs of the entire enterprise. An EDW adheres to an enterprise data model to ensure consistency of decision-support data across the enterprise.
An architecture for managing information contained in multiple formats across an enterprise. See also enterprise data architecture.
Technology providing custom views into multiple databases transparently to enable applications to more easily provide integrated real-time read and write access across databases.
A structured program for managing information as a strategic asset.
A server application component architecture defined by Sun Microsystems, Inc. Used to create application objects; related content may be sent using JavaServer Page (JSP). See Jakarta Enterprise Beans (EJB).
The collection of enterprise data models, enterprise process models, and any other model addressing the entire enterprise in scope. The complete set of enterprise models is commonly called the enterprise architecture.
An enterprise-wide program that provides a structured approach for deploying and evaluating a company’s strategy in a consistent and continuous manner. It gives an organization the capability to effectively communicate strategy and ensure that business processes are aligned to support the deployment of that strategy.
Process models of the entire enterprise at the contextual and conceptual level, typically including
- a functional decomposition,
- process flow diagrams,
- business process modeling diagrams, and
- value chain analysis linking processes to data (subject areas or entities), organizations, roles, goals, existing and planned applications, and/or implementation projects and programs.
The process of producing reports using unified views of enterprise data.
A category of software tools used to produce reports; a term for what were simply known as reporting tools.
Systems that tie together many of an enterprise’s functions, including finance, manufacturing, sales, and human resources. ERP systems enable analysis of integrated data to plan production, forecast sales, and analyze product and process quality. Many organizations extend the ERP architecture through data warehousing to support more advanced reporting, analytical, and decision-support capabilities.
The process of planning, organizing, leading, and controlling the activities of an organization in order to minimize the effects of risk on its capital and earnings. ERM includes not only risks associated with accidental losses, but also financial, strategic, operational, and other risks.
A software layer that provides data between services on an event-driven basis, using standards for data transmission between the services.
Storage designed for large-scale, high-availability environments.
Any concrete or abstract thing that exists, did exist, or might exist, including associations among these things (e.g., a person, object, event, idea, or process.).
An object that is distinguishable from others and can be tangible (i.e., person) or intangible (i.e., bank account).
The presentation of a data model diagram that shows entities, attributes of entities, and relationships between (among) entities, hence EAR.
A person, place, or thing of interest to an organization. It could be a tangible or intangible concept. May be represented by a data entity in a data model. See also entity type.
Discrete occurrences that are noted by time stamps or other ordering attributes.
The process of scanning unstructured documents to find identifiable entities, based on contextual clues.
The set of connected parent–child relationships of which an entity is a connected part. See also hierarchy.
Generally, the existence of a thing or the happening of an event.
In data modeling, a single specimen or member of an entity type population. See also object.
The changes to the entity occurrence over its lifecycle.
An entity that is at the top of a hierarchy; the basic high-level entity.
The phases and distinct states through which an entity moves through time. A state transition diagram documents the entity lifecycle.
An entity that classifies something else, or that something else refers to for clarity.
Any relationship or connection between two entities, concepts, or objects.
The graphical diagram for an entity-relationship data model. The underlying data model generally includes more semantics than is or can be represented in the view shown on the diagram (e.g., some business rules).
Generally, a record-based data modeling scheme that focuses on entities and relationships in the presentation of data model diagrams, thus suppressing the display of attributes. A true ER model allows multi-valued data items and repeating groups of items (nested relations, thus violating first normal form), and retains many-to-many relationships, attributed relationships, subtypes/supertypes, and ternary and higher-order relationships, none of which can be represented directly in a relational data model. A true ER model generally excludes (defers) the representation of entity identifiers and foreign keys. Originally proposed and named by Peter Chen in 1976.
In relational modeling, the most popular style of data model, defining entities and the business relationships between the entities. Some more detailed models include some of the attributes of these entities as well, usually those involved in the relationships as keys.
A population of entity instances that conform to the same data definition or schema, often synonymous with object type or class. An entity type represents a class of objects in the users’ universe of discourse, their world represented in a data model. They may be persons, places, things, abstract concepts, events, or other objects of interest to the enterprise.
The measurement of uncertainty in an outcome, or randomness in a system.
In the computer technology context, the conditions surrounding data, such as databases, data formats, servers, network, and any other components that affect the data.
In a business context, the influencing factors on business performance.
An aspect of an organization and its business processes defined in the DAMA-DMBOK Functional Framework. The seven environmental elements are Goals and Principles, Activities, Deliverables, Roles and Responsibilities, Practices and Techniques, Technology, and Organization and Culture.
In machine learning, one complete pass through the entire training dataset.
A relationship in which each side implies or replaces the other; the state of being interchangeable.
The study of how technology affects the health of the human body. Also known as biotechnology.
An incorrectly stated, inaccurate, or no longer valid fact.
An incorrect action taken in a process, usually resulting in a defect.
The frequency with which errors occur in transactions. Also called the failure rate.
In data quality, the percentage of data that is incorrect, inaccurate, or no longer true. Also called the data defect rate.
In artificial intelligence, the frequency with which a model’s predictions are incorrect, often used as a performance metric.
The particular value yielded by an estimator or an estimate process in a given set of circumstances.
A local area network protocol developed by Xerox in cooperation with Digital Equipment Corporation (DEC) and Intel. Ethernet uses a bus topology and supports transfer rates of 10 Mbps. The Ethernet specification served as the basis for the IEEE 802.3 standard, which specifies the physical and lower software layers.
The practice of designing and implementing artificial intelligence (AI) systems in ways that align with ethical principles, such as fairness, accountability, transparency, and privacy.
In general, a social system’s rules of behavior with which all members of that social system are expected to comply. Contrast with morals. See also professional ethics.
The occurrence of some action of interest to the enterprise, usually characterized at a point in time. For a period of time, recognizing that a process may span a duration of time, the start and stop of the process would be the events. See also transaction.
A process of analyzing notifications and taking action based on the notification content.
Data about business events (often system transactions) that have historic significance or are needed for analysis by other systems. Event data is atomic data that may be aggregated.
A type of flowchart used for business process modeling, enterprise resource planning, and business process improvement. Consists of a sequence of events in a process; the functions that execute following that event; the inputs, supporting systems, outputs, and organization units supporting that function; the control flows between events and functions; and logical decision points (branch/merge, fork/join, and OR).
Characteristic of a relationship that expresses “at most one.”
A role held by a senior manager sitting on the data governance council accountable for data quality and data practices within a department, planning and oversight of data management programs, and appointment of other data stewards. Sometimes referred to as a strategic data steward.
Business intelligence software products providing sets of reports (“briefing books”) to top-level executives. EIS offer strong reporting and drill-down capabilities, along with ad hoc query against a multidimensional database, and most offer analytical applications along functional lines such as sales or financial analysis.
An artificial intelligence system driven by rules based on the skills and experience of one or more experts in a given field, so the system processes information the same way an expert person does. Expert systems are deterministic; compare to neural networks, which are nondeterministic.
Describes a formal expression of knowledge.
Unauthorized access to data in files or applications. An exploit takes advantage of a vulnerability in a computer system to gain control and cause unintended or malicious behavior.
The process of analyzing data to suggest hypotheses using statistical tools, which can then be tested.
An entity-relationship model developed by Toby Teorey that includes more information, such as for ternary relationships and supertype/subtype relationships.
The ability to easily add new functionality to existing services without major software rewrites or without redefining the basic architecture. Also called extendibility.
Defined by a specific and finite list of values, not by conformity to any rule or requirement. Opposite of intentional.
The process of monitoring the flow of data between data sites in different organizations.
A data integration process in which data is first extracted from source systems, then loaded into a target system (such as a data warehouse), and finally transformed into the desired format within that target system. This approach is commonly used in modern cloud-based data platforms, allowing for scalable and efficient data processing.
Generally, an approach to data integration from multiple source databases to integrated target databases (operational data stores, data warehouses, or data marts).
Commonly, a software product or tool that extracts data from a data source, converts data to a new format, and loads the data to a target database. See also data integration.
An internal network or intranet opened to selected business partners. Suppliers, distributors, and other authorized users can connect to a company’s network over the internet or through private networks.
An updated Agile methodology approach to rapid application development using object-oriented techniques and a minimum of specifications. From Extreme Programming.
Describes a property that is nonspecific and unessential to a thing or event.
A verifiably true data point.
In dimensional modeling, an attribute that can be measured.
A data model that is built up from elementary facts. See object–role model.
Used primarily in Europe, the term is considered more generic. See also elementary fact sentence.
In dimensional modeling, a central table that contains numerical measures and keys relating facts to dimension tables. Fact tables contain data that describes specific events or transactions (such as bank transactions) or results from mathematical functions applied to the events or transactions (such as the net summary of a day’s transactions against a single account).
A fact table that contains data from multiple events on one row in order to track progression of steps within a process.
A fact table that contains only events with no other inherent measurements. Some have a count variable set to 1, which allows summing of these events using more optimal aggregation functions. An example is a fact table that records the attendance of a student in a class. Also called degenerate fact table.
A fact table that contains data showing the state of something at a point in time.
A fact table that contains data showing a change of some kind. Common transaction fact tables include data on sales, assignment changes, etc.
A structured process, named for Michael Fagan, that evaluates activities or operations that have entry and exit criteria for compliance with those criteria. Also called Fagan’s domain key criterion.
The frequency with which errors occur in transactions. See also defect rate.
The extent to which errors and recoveries within a distributed system are invisible to users and applications.
An incorrect result that fails to detect a condition or return a result that is actually present.
An incorrect result that detects a condition or returns a result that is not actually present.
Reducing the number of resources needed to describe a large set of data by creating new combinations of variables.
A numeric code that identifies US states, districts, and protectorates. The US government has developed various FIPS specifications to standardize such topics as weather conditions and emergency indications.
A set of databases that are documented and then interconnected to operate as one database, even when those databases are on different platforms. A data consumer gets data from the federation without knowing where the data resides.
A machine learning approach in which models are trained across multiple decentralized devices holding local data samples, without exchanging them.
The physical container for values of an attribute.
A collection of information, either on paper or electronic in the form of data fields (or more complex structures) which describe a set of entities possessing some common characteristics or attributes.
A collection of zero or more records which may have an arbitrarily complex structure (flat, hierarchical, etc.).
A Microsoft file system in which a table on a disk catalogs the location and size of files on disk, as well as free and unusable areas.
A saved set of selective criteria specifying a subset of data in a database.
A private organization whose mission is to “to establish and improve standards of financial accounting and reporting that foster financial reporting by nongovernmental entities that provides decision-useful information to investors and other users of financial reports.” FASB publishes the Generally Accepted Accounting Principles (GAAP).
The process of combining and aggregating data from different financial systems to create integrated financial analytic views and comprehensive financial statements compliant with accounting and financial reporting standards.
All financial information, including what may be “shareholder” or “insider” data that has not yet been reported publicly. Often protected by laws that restrict who manages financial data, thus assuring data integrity.
In artificial intelligence, the process of adjusting a pretrained language model on a specific dataset to enhance its performance or adapt it for a particular task.
A combination of specialized hardware and software set up to monitor traffic between an internal network and an external network (i.e., the internet). Its primary purpose is security, and it is designed to keep unauthorized outsiders from tampering with or accessing information on a networked computer system.
A method of posting a transaction in “first in, first out” order. In other words, transactions are posted in the same order that the data producer entered them.
An essential part of most artificial intelligence systems, particularly optimization-related ones. It depicts how close the system is getting to the desired outcome and helps adjust its course accordingly.
A file in which all the attribute fields are atomic, that is, single valued. See table. See also database, relational.
In a hierarchical data structure, to absorb all child records into their parent records (flattening up) or copying a parent record into each of its children (flattening down). In flattening up, each child type must be given a different name in the parent record, so that the parent record becomes a flat file with atomic fields. For example, in a hierarchical structure there might be a nested repeating group called address with a type attribute on each address instance (i.e., home, school, vacation, summer cabin). When flattened, each of the repeating attributes must be named differently, such as home street, summer street, etc. This is a technique to convert a hierarchical structure into a single flat file or relation that can be implemented more easily in a relational database management system.
A system for using significant digits and exponents to represent numbers too large or small to display using the existing format or display criteria.
From “friend of a friend.” A descriptive vocabulary using the Resource Description Framework (RDF) and OWL (Web Ontology Language). Computers use these FOAF profiles to describe social networks, including people, activities, and the relationships between them. For example, a FOAF profile might be used to find all people living in Illinois, or to list people both you and a friend of yours have in common, by defining relationships between people. Each profile has a unique identifier used to define relationships.
A system of classification that originates from collaboration of users to categorize (tag) and organize information; a usage-generated taxonomy.
A graphical symbol used to represent many-ness in the multiplicity characteristic of a relationship, preferred because it visually and intuitively communicates many-ness. First proposed by Gordon Everest in a 1976 paper. Also called inverted arrow, chicken feet, crow’s foot, or trident.
The specifications for layout or display of information, such as in a document or on a disk.
To apply display or configuration specifications to a document or dataset.
A DCMI element in element set Instantiation: physical characteristics of a resource. See also Dublin Core Metadata Initiative (DCMI).
A set of rules or instructions that, when applied to a specific input, will provide an expected output.
The process of generating physical structures from concepts and logical descriptions. See also reverse engineering.
A base artificial intelligence model trained on extensive data to mimic human language and reasoning.
A high-level programming language for manipulating a database, characterized by operations on sets of instances, that is, multiple records at a time. Contrast with low-level, one-record-at-a-time languages such as COBOL and Fortran, which are third-generation languages.
Generally, a basic skeletal structure.
Conceptually, a classification scheme used to better understand a topic; a defined and documented paradigm used as a lens to view a complex problem.
In software development, a reusable object-oriented design, including a library of reusable classes and other components, along with standards for designing additional components and how they interact.
The process of detecting patterns, trends, or correlations in consumer or corporate behavior that might indicate that fraudulent activity is taking place (e.g., identifying potential or existing fraud through analysis and comparison of standard and aberrant behaviors). See also data mining.
A tabulation of values output from a function given a set of inputs.
A list of questions and answers that are commonly seen regarding a topic.
A service that allows for the transfer of files to and from other computers. Anyone who has access to FTP can transfer publicly available files to their computer. Also called file transfer program.
Describes a system that allows communication between two endpoints simultaneously. See also half duplex; simplex.
See join, outer.
The process of searching a table space or storage file rather than using the actual table structure.
Generally, the acts, operations, and duties expected of a person or thing.
In process design, a high-level process consisting of a group of closely related lower-level activities that together contribute to the overall purpose and health of an organization or person.
In mathematics, a transformation that operates on one or more independent variables, and produces a value for the dependent variable. Generally written as D = Fn(I). See also dependency, functional.
A unit of measurement expressing business functionality provided to a user by an information system, calculated using data from past projects.
A technique of decomposing words into component parts and comparing the parts to find an acceptable level of correspondence.
An assessment of a system in comparison with another system or a set of requirements, listing those items that are not common between them.
A software product that allows SQL-based applications to access relational and nonrelational data sources.
An organized set of bilateral exchanges in which several data and metadata sending organizations or individuals agree to exchange the collected information with each other in a single, known format, and according to a single, known process.
A method for identifying and investigating configurations or relationships that affect a given problem by identifying all possible combinations and then removing those that are impossible or inapplicable.
The process of recognizing commonalities and combining similar types of entities or objects into a less-specialized type based on common attributes and behaviors, creating a supertype for two or more specialized subtypes. Contrast with specialization.
In artificial intelligence, a model’s ability to perform well on unseen data.
The process of evaluating attributes in multiple related entities for commonalities and possibly moving specialized attributes from a child or subtype entity to a parent or supertype entity where the specialization applies to more than one of the children.
The process of evaluating multiple relationships between entities in a set into fewer relationships. Usually necessary after other generalization activities have taken place and carried the relationships of the specialized entities into the generalized entities. For example, two one-to-many (1:M) relationships between two entities, each having a different parent, can be generalized into a many-to-many (M:N) relationship.
The standard framework of guidelines for financial accounting in the United States; the standards, conventions, and rules accountants follow in recording and summarizing transactions, and in the preparation of financial statements, including definitions of what the categories in financial statements mean and what practices are commonly allowed. Managed by the Financial Accounting Standards Board (FASB).
A class of machine learning frameworks in which two neural networks (a generator and a discriminator) compete against each other to generate new, synthetic instances of data that resemble the training data.
A type of artificial intelligence that can create new and original content that looks human-made, such as text, images, audio, video, or code, by learning patterns and structures from large amounts of existing data.
A popular transformer-based language model developed by OpenAI.
The science of measuring the shape of a planet and geographical point placement. This science makes global positioning systems possible. Also called geodetics.
Data used in navigation and surveying to translate positions to a position on a planet.
A system or tool for creating, storing, editing, analyzing, and managing geospatial data entities and associated attributes.
An XML grammar used to express geographical features and attributes.
The discipline of gathering, storing, processing, and delivering geospatial data.
The use of geospatial technology to survey and capture geospatial measurements.
Pertaining to data about locations on, in, above, or below a planet’s surface.
Data pertaining to locations and regions on the earth, generally expressed as latitude and longitude (and sometimes altitude); it can be located and reasoned about in terms of area.
A satellite navigation system in which satellites broadcast precise timing signals by radio to GPS receivers, allowing them to accurately determine their location (longitude, latitude, and altitude) in any weather, day or night, anywhere on Earth. GPS has become a vital global utility, indispensable for modern navigation, as well as an important tool for mapmaking and land surveying. GPS also provides an extremely precise time reference, required for telecommunications and some scientific research. The US Department of Defense developed the system, which is now operated by the United States Space Force. However, GPS is available for free use in civilian applications worldwide for the public good. See also geospatial data; geographic information system.
A special type of identifier used in software applications to provide a unique reference number. The value is represented as a 32-character hexadecimal string, and usually stored as a 128 bit integer. Ideally, a GUID will never be generated twice by any computer or group of computers in existence. The term GUID usually, but not always, refers to Microsoft’s implementation of the universally unique identifier standard.
Generally, a dictionary covering a limited subject area.
In metadata management, an extract of business metadata (terms and their meanings) from a metadata repository.
A desired state or statement of general direction for long-term improvement. See also objective.
The directional business goals of each function and the fundamental principles that guide performance of each function. One of the DAMA-DMBOK Functional Framework Environmental Elements.
A dataset with perfect contents according to requirements. Used to describe sources of record or datasets used for testing purposes. Also called gold standard; gold source. See also system of record.
Generally, the exercise of authority and control over a process, organization, or geopolitical area.
In data management, the process of setting, controlling, administering, and monitoring conformance with policy. See also data governance.
An optimization algorithm used to minimize the loss function in machine learning models.
The degree of summarization of data in a dataset, representing the most detailed information available. All aggregations would result from mathematical functions using this dataset, without any additional data. Also called granularity.
Generally, a set of homogeneous nodes (vertices) and edges (arcs) between pairs of nodes.
In business intelligence, a visual representation using references to a set of axes to illustrate the relationship between functions or sets of quantities. See also chart.
A type of neural network designed to process data structured as graphs.
The study of mathematical structures used to model relations between items within a dataset.
A system that allows users to interact with a computer through visual elements such as icons and pointers, rather than text only. The dominant style for desktop and client/server computer applications.
A processing unit built to perform large numbers of calculations in parallel, which makes it especially useful for graphics and artificial intelligence.
Internationally accepted civil calendar used in the western world, with additional rules regarding application of leap days and other minor adjustments.
A web-based operation allowing companies to share computing resources on demand.
A distributed data storage technology developed by Apache Software Foundation.
Describes a system that allows communication between two endpoints where only one may transmit at a time. See also simplex; full duplex.
A class of binary linear codes used for parity calculation that can detect up to two simultaneous bit errors, rather than just odd numbers of errors. Named for Richard Hamming. Used in computer random-access memory (RAM) and telecommunications for validating data transmission.
Data allocated in an algorithmically randomized fashion in an attempt to evenly distribute data and smooth access patterns.
To calculate a hash key for data.
A law enacted by the US Congress in 1996. Title I of HIPAA protects health insurance coverage for workers and their families when they change or lose their jobs. Title II requires the establishment of national standards for electronic health care transactions and national identifiers for providers, health insurance plans, and employers. Title II also addresses the security and privacy of health data. The standards are meant to encourage use of electronic data interchange in US healthcare.
A runtime memory area used to dynamically allocate and manage objects and data during program execution.
Made up of dissimilar or varied elements.
Data originating from different systems, formats, or technologies. Modeling such data involves schema mapping and integration.
Approximation methods and “rules of thumb” for obtaining a goal, a high-quality solution, or improved performance. It sacrifices completeness to increase efficiency, as some potential solutions would not be practicable or acceptable due to their rareness or complexity. This method may not always find the best solution, but it will find an acceptable solution within a reasonable timeframe for problems that would require almost infinite or longer than acceptable times to compute.
A numbering system using a base of 16 and using letters A through F to represent 10 through 15 decimal. A byte is generally 8 binary digits, so that 1 hexadecimal representation represents 4 binary digits. Core dumps are expressed in hexadecimal, for example.
A high-level language on a hierarchical data structure, similar to SQL on a multi-file data structure.
Generally, a classification structure arranged in levels of detail from the broadest to the most detailed level. Each level of the classification is defined in terms of the categories at the next-lower level of the classification.
In dimensional modeling and dimensional databases, the organization of a dimension’s members based on parent–child relationships, typically where a parent member represents the consolidation of child members.
A column with a large number of unique values, such as user IDs or email addresses. Important for indexing and performance tuning in data models.
Optional guidance given to a system or user to influence behavior or performance without enforcing it.
A chart that shows quantities of data points that occur within various numeric ranges.
The reinterpretation of historical data based on new data, validated or invalidated assumptions about the data, or different perspectives on the environment that generated the data.
Software that facilitates reading, writing, and managing large datasets residing in distributed storage. An open-source project of Apache Software Foundation.
Made up of elements that are all the same or similar nature. Opposite of heterogeneous.
Describes multiple members in a set that have no differences in nature or structure.
A set of nodes conforming to the same definition (or of the same type).
A term that has the same or nearly same spelling or sound as another term, but has a different meaning. Contrast with synonym.
A network-connected computer or system that provides or runs services, applications, or resources for other systems.
Describes a processing method in which the host computer controls the session. A host-driven session typically includes terminal emulation, front-ending, or client/server types of connections. The host determines what is displayed on the desktop, receives user input from the desktop, and determines how the application responds to the input.
Consolidating related names and addresses into groups.
A system of communicating between a web browser application and a document stored on a web server.
In data vault modeling, a core entity table that contains the unique business keys and metadata such as load date and record source connecting them to satellites and links.
An interface from a system to a human that enables the human to interact with and receive information from that system. See also interface.
The application of advanced technologies, including artificial intelligence and machine learning, to increasingly automate processes and augment human efforts.
An OLAP product that stores all data in a single data cube which has all the application dimensions applied to it.
A one-way reference from one electronic document to another. Most frequently implemented as navigational links from one web page to another.
Electronically stored text data organized into documents and logical sections that can be accessed randomly via hyperlinks as well as sequentially.
A specific style of notation for business process flow diagrams in process models. See also Integrated Definition (IDEF) Methods.
See data modeling notation, IDEF1X.
A style of notation for modeling resource behavior over time in a simulation. See also Integrated Definition (IDEF) Methods.
The label (value, name, handle, etc.) used to unambiguously refer to individual instances of a population. It is represented by a key in the records or relations of a database. There must be a one-to-one relationship between the values of the key and the members of the population in the user world. Identifiers define keys (primary keys, candidate keys).
A class word assigned to attributes or columns containing unique identity values for that instance or row.
A DCMI element in element set Instantiation: a unique reference to a resource. See also Dublin Core Metadata Initiative (DCMI).
The process of identifying and detecting an object or feature in a digital image or video.
The quality that once data has been written, it cannot be changed. This ensures the permanence of the data and trust in its provenance.
Identifying the potential consequences of changing an object to its related objects.
Installing and converting to use of a software application.
Having disagreement or disparity among things or parts of things. Having internal contradictions.
See phased implementation.
Data propagation to a target database limited to the data that has changed in the source database since the last load.
Generally, a cross-reference created to find something that matches some selection criteria.
In data management, a data structure that cross-references a set of values from the same domain to the places (records or rows) where each value appears, generally within a single file (see join index). An index is usually ordered according to the values in the domain. In general, an index can have multiple references (or pointers) for each value, unless the index is on an identifier, in which case there is a one-to-one relationship between the values and the record identifiers. An index is used to improve retrieval performance on a file; it does not add any new information to the database.
To create a cross-reference list.
An indexing technique in which a separate structure stores the references to the data as bit arrays.
Describes an index in which every key relates to a block in a data file, using the lowest search key in the block.
A binary search tree index that stores index pointers in block partitions according to the values themselves. It simulates a binary search tree and uses corresponding search methods to give performance of the order of Log(base2)N, rather than N as in conventional indexes.
An indexing technique in which the actual data is physically stored in the order of the index values, rather than having the index in a separate structure pointing to the data rows. Only one clustered index may exist on an object at a time.
An index in which the values of the data are stored in the index, allowing data retrieval from the index itself instead of the data object.
An index in which every row in the indexed structure relates to a value in the index.
A type of index that is either related to a non-partitioned table or not partitioned even though the underlying table is partitioned.
An index structure that stores locations of keywords within a set of files, and possibly the location within the file, rather than a list of possible values, in order to provide speedy searches for words or phrases. Mostly used for content searches through multiple files, such as a search for the term DAMA within several web pages or documents.
A type of partitioned index in which the index block corresponds to one and only one data block.
An index in which the actual data is stored in random order, not physically in the order of the index. Files can have multiple non-clustered indexes, and each non-clustered index will take up space as an object.
An index in which the value being indexed is reversed (reversing the characters or reversing the digits) before being sorted. This is especially useful for indexing sequence numbers, where the most significant digit rarely changes, but the least significant digit always does.
An index in which every possible value in the indexed object relates to a pointer in the index, and few of those values actually appear in the indexed file or object, so that the index is mostly empty. See also index, block.
An index on an identifier, or attribute(s) defined as unique, in which case there can only be one pointer for each value entry in the index.
A method of storing data such that the index key controls the physical order of the data within the file.
A form of disk file management that uses indexes to assign storage information and retrieve data from the disk. Originally the name International Business Machines Corporation (IBM) assigned to its partial or block indexing scheme. Has since taken on a more generic usage. See also index, block.
An attribute type that is considered to be binary: On or Off, True or False, Yes or No.
A class word, abbreviated usually to ind.
In data management, the process of creating categories from instances.
In artificial intelligence, the assumptions a learning algorithm makes to predict outputs for unseen inputs.
Reasoning from known propositions.
The process where a trained artificial intelligence model makes predictions or decisions based on new input data. Essentially the “thinking” part after the model has been trained.
A model in which some of the data is inferred from actual data points.
A visual representation of data or information intended to present details quickly and clearly. Meant to enhance the human visual system’s ability to see patterns and trends.
Generally, understanding concerning any objects, such as facts, events, things, processes, or ideas, including concepts that, within a certain context and timeframe, have a particular meaning.
The interpretation of data based on its context, including
- the business meaning of data elements and related terms,
- the format in which the data is presented,
- the timeframe represented by the data, and
- the relevance of the data to a given usage.
Synonymous with information technology, used predominantly in the United Kingdom, particularly in the UK education system.
Data in any form, or media placed into meaningful context for users, collected in relation to business or research activity.
Formal management of data, and organization of the users of that data, that provides context and creates information assets.
Established by the US DARPA to collect and integrate personal information of US citizens and residents, primarily targeted for data mining to detect threats to the country.
Confusion in information that may be relevant and timely, but is interpreted incorrectly, inconsistently, or incompletely.
A person or group that receives data and uses it to create information. A more descriptive term for a data consumer, since the consumer creates and uses information by interpreting data in context.
A collection of the metadata that relates to data warehouse and business intelligence systems within an organization, providing some context to the metadata to make it usable and searchable by business professionals in natural language terms. The directory includes business metadata, such as definitions, domains, examples, relationships, functions, rules, advisories, and equivalents in other environments. It also may include technical metadata about datatypes, lengths, number of distinct values, transformation rules, and replication schedules.
And engineering discipline that deals with the generation, distribution, analysis and use of information, data, and knowledge across various systems. Developed by Clive Finkelstein in the 1970s and popularized by James Martin; incorporates a record-based data modeling scheme and notation convention.
See data modeling notation, Information Engineering (IE).
Depicts the complete flow of information from source to target.
In data warehousing, shows the flow of data from all sources through intermediate structures into final targets where the data is turned into information.
An approach to manage the flow of a system’s information from creation through usage to purge.
The management of data in context, with relevance and timeframes, for business benefit.
A technique of dividing and categorizing information for ease of comprehension and recall.
A model showing information structure, usually at a conceptual or logical level.
The identification and study of the information needs required to satisfy a particular business driver.
The state in which the rate or amount of input to a system or person outstrips the capacity or speed of processing that input successfully.
A statement of principles and guidelines for information management.
The degree to which information, as prepared from the data, meets the requirements and expectations for that use.
The rate at which information loses relevance over time if not refreshed and reviewed.
A form of data quality management with an added emphasis on managing the quality of the context in which data appears as well as the quality of the data itself.
The full set of data and processes – technical, procedural, and organizational – that collect,
transform, and distribute information appropriately.
Generally, an automated or manual organized process for collecting, manipulating, transmitting, and disseminating information. See also application.
In data management, a system that supports decision-making concerning some piece of reality (the object system) by giving decision-makers access to information about relevant aspects of the object system and its environment.
The first phase in an Information Engineering methodology. The goal of ISP is to define an enterprise architecture. ISP is usually performed as a separate project, defining several subsequent projects. See also business systems planning.
A broad subject concerned with technology and other aspects of managing and processing information, especially in large organizations. IT deals with the use of electronic computers and computer software to convert, store, protect, process, transmit, and retrieve information.
The department of an organization that deals with computer hardware, application software systems, and data. See also management information system (MIS).
The process of making decisions about information technology (IT) investments, the IT application portfolio, and the IT project portfolio.
A framework of supplier-independent best practice management procedures for delivery of high-quality IT services.
The budgeting, funding, issue and risk management, and overall tracking mechanism for all information technology (IT) projects and programs.
The formal process for managing IT assets, including application software, infrastructure software and hardware, internal staff, and external consulting, and how they support business processes and strategies, outside of program or project management.
The governing body of senior executives responsible for aligning information technology (IT) goals, objectives, strategy, architecture, and projects with enterprise goals, objectives, and strategy, for oversight of IT functions and projects, including project prioritization and funding.
A process to link conceptual and logical data models to process models, applications, organizations, roles, and/or goals, to provide context, relevance, and timeframes.
An approach to data warehousing that supports the implementation of central, functional, or decentralized warehouses. It may provide information, but it does not contain information by itself. See also data warehouse.
The underlying foundation of a system or organization. See also infrastructure, information technology (IT).
A combination of technologies and the interaction of technologies that support a data warehousing environment.
The complete set of hardware, operating system, and software products implemented in support of the application software of an enterprise.
The infrastructure organization responsible for design, implementation, maintenance, operation, and support of the information technology infrastructure.
Generally, something received by succession.
In data modeling, the sharing of the attributes and behaviors of parent class (supertype entity).
A SQL statement (command) that specifies addition of rows of data in a relational database.
Moving a software product or application into a production computing environment.
An individual member of a population, such as a value in the domain of values for an attribute, or an individual entity record in a file. See also entity instance; attribute; object.
A set of facts describing an actual entity occurrence at a point in time or during a period of time. The data about an occurrence may vary in different instances.
An instance of a software object or database row/record.
The name of a DCMI element set (Date, Format, Identifier, Language). See also Dublin Core Metadata Initiative (DCMI).
A nonprofit consortium of professional associations with the common goal of assessing, credentialing, and improving the skills and standards of students and individuals employed in the business, computer, information, and communications technology industries.
A professional organization for engineers, including software engineers.
A set of rules or other formal set of instructions assigning responsibility as well as the authority to an organization for the collection, processing, and dissemination of information.
An automated method of gathering information through measurement devices such as cameras, smartphones, health monitors, or utility meters. Often processed immediately without storage because of the delivery rate of speed.
A natural whole number (positive or negative) or zero. From the Latin integer, for “intact, untouched.” Contrast with real number.
To form or blend into a whole; to unite with something else; to incorporate into a larger unit; to bring into common organization.
A data resource that is fully integrated within a single, organization-wide, common data architecture and is deployed as necessary to meet the business information demand.
ICAMS (Integrated Computer-Aided Manufacturing) Definition Languages, developed for the US Air Force. There are several of these modeling languages.
- IDEF0 describes functional modeling notation.
- IDEF1X describes data modeling notation.
- IDEF2 describes simulation model notation.
- IDEF3 describes process description capture.
- IDEF4 describes object-oriented design.
- IDEF5 describes ontology description capture.
- IDEF6 describes design rationale capture.
- IDEF7 describes information system auditing (not developed).
- IDEF8 describes user interface modeling.
- IDEF9 describes business constraint discovery.
- IDEF10 describes implementation architecture modeling (not developed).
- IDEF11 describes information artifact modeling (not developed).
- IDEF12 describes organization modeling (not developed).
- IDEF13 describes three schema mapping design (not developed).
- IDEF14 describes network design.
A software application or suite of integrated applications used to design, develop, and test application code.
A disk system where the disk controller is integrated into the drive itself, rather than remotely.
The unified state of multiple components in one whole, complex system.
The process of unifying multiple components into one complex system.
A rule that ensures the accuracy and consistency of data (e.g., primary key, foreign key, a unique constraint).
Intangible assets of an enterprise created by its knowledge workers, including information about tangible assets, documents, ideas, patents, inventions, trade secrets, brands, software, and databases, and generally expressed in some copyable storage form.
The name of a DCMI element set (Contributor, Creator, Publisher, Rights). See also Dublin Core Metadata Initiative (DCMI).
The ability to understand and apply to practice.
In common use, a collection of data about something or someone.
A software routine that waits in the background and performs an action when a specified event occurs. For example, agents could transmit a summary file on the first day of the month or monitor incoming data and alert the user when certain transactions have arrived.
In artificial intelligence, an autonomous entity that perceives its environment and takes actions to achieve specific goals, such as robots or software agents that act with purpose.
Describes a set of valid values defined by conformity to rules. Each time the rules are executed, the result set may be different from the time before. For instance, the set of customers with overdue balances is an intensional set. See also domain; extensional; master data management (MDM).
A set where membership is defined by explicit rule(s) applied to members of a larger set. The operands of the rule would be attributes of the entity instance being considered for membership. Opposite of extensional set.
A query formed through the interaction between a human and the (computer) system. The system can assist the user in formulating a query. The query may then be executed (as it usually is) or stored for later execution.
The degree to which attributes in a set influence each other’s values.
In data quality, the degree to which one attribute or row influences the values of other attributes or rows.
The connection to and means of communication between people and systems, or between different systems.
The standard application programming interface for calling Common Object Request Broker Architecture services.
The international standards organization that determines Generally Accepted Accounting Principles (GAAP).
A global network that identifies what international standards are required by business, government, and society; develops them in partnership with the sectors that will put them to use; and delivers them to be implemented worldwide. ISO is the world’s leading developer of international standards.
A unique commercial book identification number, based on the Standard Book Numbering code, that is applied to books and book-like products that are published internationally.
A branch of the United Nations that sets and manages the representation of phone numbers (among other things in the industry). Formerly the International Telegraph and Telephone Consultative Committee (commonly abbreviated CCITT, from French: Comité Consultatif International Téléphonique et Télégraphique).
The global set of computers linked over public networks and addressing each other through data source names (DSNs) and URL addresses, using HTTP (Hypertext Transfer Protocol) for their primary access protocol and HTML (Hypertext Markup Language) to display information.
A nonprofit digital library offering free access to uploaded books, music, and archived web pages.
The address to an internet site that has been saved with a name or a tag.
A network of physical objects or “things” embedded with software, sensors, and other technologies to connect and exchange data with other devices or systems over the internet.
A unique identifier for each computer or other device on a network including the internet. This string of numbers allows computers, routers, printers, and other devices to recognize or identify one another and communicate.
The process of adding attributes to sites on the internet in order to enable grouping or filtering.
The ability of various types of computers and programs to work together and share data across different platforms.
The use of a formula to estimate an intermediate data value.
A computer language that compiles source instructions one at a time as needed at run time.
Generally, a question; a sentence that generates a reply.
In language, a part of speech that is used to show a question: Who, what, when, where, how, and why are all interrogatives.
Of or relating to questions.
A SQL set operator that intersects two tabular SELECT answer sets with consistent column structures into one answer set table in which only rows that match using the join conditions are included.
An entity used to represent a many-to-many relationship between two other entities. See also data entity, associative.
A numeric scale in which the numbers have no arithmetic zero point or origin. Thus, it is only meaningful to add and subtract them, not multiply or divide. We cannot say that 60 degrees is twice as hot as 30 degrees. Examples are date, time, and temperature, except for Kelvin, which does have a meaningful absolute zero.
A subset of the internet used internally by an organization; the use of internet technologies over a private network. Unlike the larger internet, intranets are private and accessible only from within the organization.
Describes a property that is specific and essential to, and inseparable from, only one thing or event, and that is independent of any other property.
The portion of total web content that consists of material that is not accessible by standard search engines. It is usually to be found embedded within secure sites, or it consists of archived material.
International standards for quality management, specifying guidelines and procedures for documenting and managing business processes and providing a system for third-party certification to verify those procedures are followed in actual practice.
A Java application programming interface for building of enterprise software. Formerly Enterprise JavaBeans.
A cross-platform, object-oriented programming language that allows applications to be distributed over networks and the internet.
A standard application programming interface for accessing relational data from Java programs.
The standard application programming interface for sending and receiving messages for Java programs.
A collection of technologies used to create dynamic web content. Also used to generate and consume XML between n-tier servers or between servers and clients.
A standard for developing multi-tier applications, particularly for middleware and application servers. In 2019 the platform was renamed Jakarta Enterprise Edition.
A series of scripts or programs that run at a predefined schedule without manual intervention for the manipulation, movement, transformation, archiving, or backing up of a set of data.
In relational databases, an operation in which the data from two sets is combined into a larger result set based on common or matching values in each set.
A form of table join where rows from the table on the left side of the join conditions are returned, regardless of whether there is a match in the other table. The SQL query clause “WHERE A.JC = B.JC” returns all rows in A table plus rows in B table where B’s join conditions match A’s join conditions.
A join in which some entries are included in the join which do not appear in both of the joined files. If all entries in each of the two joined files are included, it is called a full outer join, or simply outer join. Also called partial outer join.
A form of table join where rows from the table on the right side of the join conditions are returned, regardless of whether there is a match in the other table. The SQL query clause of “WHERE A.JC = B.JC” returns all rows in B table plus rows in A table where A’s join conditions match B’s join conditions.
A group process for defining requirements and designing a computer-based system. JAD sessions are highly focused, bringing together business professionals and information technology professionals under the leadership of a skilled facilitator. JAD sessions can be used to draft, review, and refine data models. They are generally held over one or more contiguous days. JAD sessions are more valuable as brainstorming sessions, getting a sense of the attendees, taking straw polls, and so on.
A modeling technique in which some attributes describe a class in conjunction with other attributes that also describe that class. Named for Thomas Bayes. See also predictive modeling.
Generally, a written record of observations and experiences.
In data management, a file that contains database activity details for rollback and recovery. See also log.
A standard file format for compression of photographic images.
A solar calendar that established months and years, with a leap day every four years. Supplanted by the Gregorian calendar in AD 1582.
The date expressed as a simple number, used by astronomers and historians due to the simple math involved. The Julian calendar started on January 1, 4713 BCE, at noon, where BCE stands for “before the Common Era.” The Julian date for noon on CE 2011 February 20 is JD 2455613.00000, where CE stands for “Common Era.”
Delivery of information at the time it will be used, not before and not after.
The Japanese word for “continuous improvement.”
A function used in algorithms, especially in support vector machines, to transform data into a higher-dimensional space.
A data item or combination of data items designated to uniquely identify a particular entity instance or table row. See also identifier (ID).
A business calculation (metric) with associated target values or ranges that allows macro-level insights into the business process to manage profitability and monitor strategic impact.
A unique identifier for an entity instance other than the primary key. Usually an alternate key is a unique natural key. Also called secondary key.
Also called domain key, natural key.
An identifier familiar to and used by data consumers, using existing attributes, which has a logical relationship to the attributes within the row.
A primary key that uses attributes that have meaning to the business. Opposite of a surrogate key.
A key that can uniquely identify occurrences of an entity. Each occurrence must have a different key value, and every attribute in the key is needed to uniquely identify each occurrence. Such identifiers are “candidates” to become a primary key, and candidate keys not selected as the primary key are considered alternate keys.
One or more attributes in a relational table that is from the same domain as the identifier of the same or another table; can be thought of as a logical pointer from the “referencing” entity table (with the foreign key) to the “referenced” entity table (with the identifier). It is used to represent a many-to-one relationship between the referencing and referenced tables. It is not necessary for a foreign key to have a value; that is determined by the independently defined dependency characteristic.
The preferred primary key of a parent data subject that is placed in a subordinate data subject to identify the relevant parent data occurrence in that parent data subject.
A set of one or more data attributes whose values are used to uniquely identify an entity instance or relational database table row. The primary key will have a unique value for each record or row in the table and is the means of navigation across entities and tables. Primary key attributes and values of parent entities and tables appear as foreign key attributes and values in child entities and tables.
A key whose value identifies a set of occurrences in a data structure that share common characteristics. Access by secondary keys may return multiple occurrences, where access by a primary key is assured to find no more than one occurrence.
A set of attributes in a dataset such that there are no repeated value sets. Each combination of the values in the attributes in a superkey is unique.
A single-part, artificially established, physical identifier for a dataset, usually not visible to business users, and used for database management and performance. Surrogate key assignment is a special case of derived data: one where the primary key is derived. A common way of deriving surrogate key values is to assign integer values sequentially. Sometimes referred to as a dummy key, sequential key, or auto-number field.
The practice of using a software program or hardware device to record all keystrokes on a computer keyboard, either as a surveillance tool or as spyware.
Used by employers to monitor employee computer habits.
A term found in a document, indexed to enable document search and location.
A modeling technique that assigns values to points based on the values of the k-nearest points, such as average value or most common value. See also predictive modeling.
Generally, understanding; familiarity or expertise gained through experience or association; cognizance, the fact or condition of knowing something; the acquaintance with or the understanding of something; the fact or condition of being aware of something, of apprehending truth or fact.
Understanding; awareness, cognizance, and the recognition of a situation and familiarity with its complexity. Understanding of the significance of information; information in perspective, integrated into a viewpoint based on the recognition of patterns (such as trends and causes) based on other information and experience.
In artificial intelligence (AI), a structured database of facts and rules used in expert systems or knowledge-based artificial intelligence for reasoning and decision-making.
A database of rules, usually expressed in an if/then format.
Knowledge that is easily codified, shared, documented, and explained.
A graphical representation of real-world entities and their interrelations, often used in natural language processing and semantic search.
A standard format for exchanging rules between artificial intelligence systems.
A collection of methods relating to a multidisciplinary approach to achieve organizational objectives by creating, organizing, and sharing information and knowledge.
In artificial intelligence, the way information is structured so that a computer can use it to solve complex tasks.
Knowledge that is based on experience and not easy to share, document, or explain.
Anyone who works for a living by understanding information. A type of information consumer. Knowledge workers seek to gain expertise through the understanding of information and then apply that expertise by making informed and aware decisions and actions.
A title or tag applied to a data attribute that concisely describes the entity or attribute type and/or content for ease of sorting, filtering, or scanning for relevance.
An area of a data warehouse system where data is first placed after extraction from a source system. See also staging area.
A system of communication using sounds (spoken language) or symbols (written language).
A DCMI element in element set Instantiation: the terminology set used to describe a resource. See also Dublin Core Metadata Initiative (DCMI).
A computational model trained on immense amounts of data, making it capable of understanding and generating natural language and other content needed to drive multiple applications and resolve multiple tasks. LLMs have helped with the development of artificial intelligence.
The last date that an attribute or entity instance is valid.
Variables in statistical models that are not directly observed but are inferred from observed data.
A group of functionally related components within an architecture representing a level of abstraction different from other layers within the architecture.
In artificial intelligence, a building block of neural networks in which computations occur.
The average time it takes a person to learn how to use or master a tool or technique.
Data that comes from production files and databases that stand outside of, or came from a previous form of, the organization’s data architecture.
An application implemented outside of, or from a prior version of, an organization’s application architecture. Typically an older application that may be slated for eventual replacement. Legacy systems are often frustrating because they are difficult to change, few people know exactly what they do and how they do it, and/or the technology on which they are dependent is becoming obsolete and unsupportable.
A group of codes characterized by homogeneous coding, and where the parent of each code in the group is at the same higher level of the hierarchy.
Taking full advantage of a resource to effectively achieve a desired outcome.
In general, a glossary or dictionary.
In data management, a computer-readable data dictionary of attributes.
Possession of and responsibility for current economic costs, such as a debt; the opposite of an asset.
In charts, a ratio of the size of a graphical representation of an item or effect to the size of the effect within the data itself. Describes how far off the graphic representation shown is in respect to the actual data driving the chart.
The set of valid states of an object, arranged in sequence from “birth” to “death.” Usually depicted in a state transition diagram.
A shorthand reference to the software development lifecycle.
Managing data throughout its life from creation and initial storage to the time it becomes obsolete and is deleted. See data lifecycle.
The present value of future cash flows expected from something (equipment or property) or someone (customer or citizen) over an anticipated timeframe, computed using the costs of acquisition and retention, in order to estimate profitability. Also called customer lifetime value.
The path that a data attribute travels between systems, and the alterations made during that journey.
The path that metadata travels between the source systems and the metadata repository.
Relating to a line, or with a progression that strongly resembles a line.
A linear approach to modeling the relationship between a dependent variable and one or more independent variables.
The process of connecting related records from different sources based on common attributes.
To place data into a target database, data warehouse, data mart, or repository. See also extract-transform-load (ETL).
To distribute workloads across multiple systems to optimize performance and reliability.
Facts and figures about a load process, such as number of records loaded, number rejected, reject reason, etc. Often stored as metadata.
A computer network covering a limited physical area, such as an office or a building.
Occurs when one process requests and is denied a lock to a resource because it is held by another process.
In data management, a collection of records that describes the sequence of events that occur during a database management system (DBMS) execution, recorded for use in database recovery in the event of a DBMS failure. See also journal.
A modeling technique where unknown values are predicted by known values of other variables where the dependent variable is binary type. See also predictive modeling.
In artificial intelligence, a statistical model commonly used for binary classification tasks in machine learning.
The management of flows of goods, information, resources, etc. in a logical progression between points of origin, consumption, and destruction.
A process of writing data in which modifications are written to a log before being applied to the stored data at rest.
A type of recurrent neural network that can capture long-range dependencies in sequences.
A data structure used to map input values to output values, often for quick reference.
In databases, a table that stores standardized reference data.
An arrangement whereby components can be easily attached and detached, enabling easier configuration changes. See also design.
Data automatically created by machines via sensors or algorithms or any other nonhuman source. Commonly known as internet of things (IoT) data.
A subset of artificial intelligence that involves the use of algorithms and statistical models to enable machines to improve their performance on a task through experience.
A standard for representation and communication of bibliographic and related information in machine-readable form, created by the US Library of Congress. From Machine-Readable Cataloging.
A stored sequence of commands or instructions that, when invoked, executes a series of commands or keypresses. Commonly used to automate repetitive tasks within applications such as word or number processors.
The point on Earth’s surface at which the magnetic field points vertically down from the northern hemisphere. Not the same as true north.
A centralized computer architecture, once dominant and still widely used, supporting a very large number of applications.
A modeling technique that includes rules that result in non-outlier data added directly into the model calculations. See also predictive modeling.
An award given by the US National Institute of Standards and Technology to recognize total quality management achievements of US business, health care, and educational organizations. It was established in 1987 and named for Malcolm Baldrige Jr., who was US secretary of commerce from 1981 until his 1987 death in a rodeo accident. The purposes of the award are to promote quality awareness, recognize quality achievements of the US companies, and publicize successful quality strategies.
An abbreviation for malicious software; software used to disrupt computer operations, gather sensitive information, or gain access to computer systems.
The possibility of something being controllable and supportable.
Describes the ability to create and maintain an productive environment.
The ability to deliver consistent, predictable access to data whenever users need it.
The operational implementation of metadata architecture, including a metadata repository, metadata sources, integration procedures, management processes, delivery procedures to metamarts, and access interfaces.
Planning for and control of replicated data, ensuring there is a master record and that copies of that record are consistent, and that minimal redundant and nonproductive replication occurs.
A reporting or business intelligence system. See also information system (IS).
Required, not optional. A dependency that must be fulfilled.
In SQL and many database management systems, equates to NOT NULL or NOT NULLABLE constraints.
The characteristic of a relationship in which a member of one population can be related to multiple members of the other population, and vice versa. Sometimes notated M-N or M-M. See also cardinality; relationship.
The reverse of one-to-zero or one-to-many.
To associate mathematically every member in a given set with at least one member of another set.
A list of source and target entities and attributes linked by a set of instructions.
The use of a fixed list of specific items to track the progress of inflation in an economy or market. This list contains a number of the most commonly bought food and household items. Variations in the prices of these items from month to month gives an indication of the overall development of price trends.
The process of identifying groups of potential customers with similar needs and/or characteristics who are likely to exhibit similar purchase behavior.
Software that helps with the up-front planning of a marketing function and the coordination and collaboration of marketing resources.
Tags and other annotations inserted to identify sections in a document.
mark up: To annotate documents by inserting tags to offset and identify sections.
A combination of application outputs, content objects, or data attributes that creates new structures from the parts.
A display of nonintegrated data attributes from multiple sources that can be combined to form new display objects.
The definition and delivery of customized products and services on a wide-scale and cost-effective basis, typically by leveraging information technology. A concept defined and developed by Joseph Pine of IBM.
In computer architecture, the “shared nothing” approach to parallel computing. A distributed-memory computer system of multiple nodes where each node has data, processors, memory, and a network link, so that each node may process a part of a task independently on its data and then send the results back to a collector. Growth is achieved by adding more nodes. Possible bottlenecks include network bandwidth. Requires specialized partitioning to spread the data effectively and efficiently across the nodes based on expected usage. Contrast with symmetric multiprocessing.
The data that provides the context for business activity data in the form of common and abstract concepts that relate to the activity. It includes the details (definitions and identifiers) of internal and external objects involved in business transactions, such as customers, products, employees, vendors, and controlled domains (code values).
The processes that control management of master data values to enable consistent, shared, contextual use across systems of the most accurate, timely, and relevant version of truth about essential business entities.
Master data about an organization’s financial configuration, including business units, cost centers, profit centers, general ledger accounts, budgets, projections, and projects.
Master data about locations specifically related to a business, in the form of geographic data, such as business party addresses and facility locations.
Master data about individuals and organizations, and the roles they play in business relationships. May include customers, employees, vendors, partners, and competitors; citizens; suspects, witnesses, and victims; members and donors; patients and providers; or students and faculty.
Master data that focuses on an organization’s internal products or services, or an entire industry’s shared products or services, including competitor products and services.
An old term for “database,” used before relational databases were commonplace. Now used as a concept in master data management regarding the official version of master data.
The process of comparing rows in datasets to determine which rows describe the same thing and are therefore either complementary or redundant. See also similarity analysis.
A view that is stored as a separate object in order to optimize performance.
A set of arrays of the same type, where each array is seen as a dimension. Matrices are used to analyze and document the linkages and relationships between the occurrences of one dimension with the occurrences of the other dimensions. See also array; scalar.
A structured collection of characteristics of effective processes at progressive levels of quality and effectiveness. A maturity model provides a common language and shared vision for process improvement, a standard for benchmarking, and a framework for prioritizing actions. A maturity model assumes a natural evolutionary path for organizational process improvement.
The result of dividing the sum of all values within a set by the count of all values included.
The predicted elapsed time (arithmetic mean) between system failures. Used to evaluate system stability.
The predicted elapsed time (arithmetic mean) between a failure and restoration of a system. Used to evaluate system support efficacy.
Loosely used, a metric.
In data modeling, a quantified characteristic; the unit used to quantify the dimensions, capacity, or amount of something.
To quantify one or more dimensions; capacity or amounts of something.
The centermost value in an ordered set of values. If the set quantity is even, then the average of the two centermost values.
A comprehensive set of descriptors used to index medical and life sciences journal articles.
All information regarding a person’s health or medical treatments. In the United States this information is restricted under HIPAA (Health Information Portability and Accountability Act).
An individual instance of a population.
The state of belonging to a set.
A system and method for placing, locating, and indexing blocks of data.
An artificial intelligence neural network architecture designed to store and retrieve information, mimicking human memory.
An electronic request or reply expressed in data. Messages can be expressed in the form of XML (Extensible Markup Language) documents.
A software intermediary function that dispatches messages to the correct sites.
Software that enables intercomponent communication through messages and message routing through a message broker.
Literally, “data about data”; data that defines and describes the characteristics of other data, used to improve both business and technical understanding of data and data-related processes.
Metadata that records lifecycle attributes of a resource, including acquisition, access rules, locations, version control/differentiations, lineage, and archival/destruction.
The names and business definitions of entities and tables, attributes and columns, and defined domain data values that establish the consistent shared meaning of data.
Nontechnical metadata of interest to business professionals, ideally defined by business data stewards. Includes the names and definitions of business entities and their data attributes in a conceptual or logical data model, as well as the equivalent business definitions for tables and columns in a physical data model or implemented database. Business metadata also includes the descriptions of business relationships between business entities, the business rules that govern those relationships, the logical business names and definitions of domain values (code values), and the descriptions of rules governing use of these code values.
Metadata about data stewards, stewardship processes, and responsibility assignments.
Metadata that characterizes and catalogs the actual resource.
The process of joining differing attributes in multiple metadata repositories to allow for easier access.
Processes that create, control, integrate, access, and analyze metadata.
Metadata that describes the physical condition of stored resources and changes to that physical condition over time (such as copying to different media).
Metadata that defines and describes the characteristics of other systems (processes, business rules, programs, jobs, tools, etc.).
Generally, any structured database of metadata, often in support of a particular tool.
Specifically, an integrated database of metadata, considered the official representation of metadata in an enterprise. Contains business and technical metadata from multiple sources. It may be updated in real time or in batch.
Metadata that represents rules regarding the use of that resource with respect to intellectual property rights.
Metadata that describes resources at atomic levels, and at higher levels, including how the atomic data attributes are related.
The process of consolidating and relating data attributes with the same or similar meaning from different systems.
The physical characteristics of data found in a database, including physical names, datatypes, lengths, precision and scale of numeric data attributes, statistics, source locations (lineage), and code values. May also include data about programs and other technology.
Metadata that represents how the resource is accessed, processed, and output.
A data store for metadata fed from a metadata repository and created for a specialized audience or tool, such as an information directory. Also called metadata mart.
Generally, a model that specifies one or more other models.
In metadata management, a model of a metadata system or a data model for a metadata repository.
Generally, a formalized system of principles, practices, and procedural methods used to build systems, perform a process, or solve a problem, including organizational arrangements, deliverables, and timelines.
In object-oriented design and programming, a function bound to a class as part of its overall behavior, executed in response to a message.
The study of methods.
Generally, a unit of measure selected used to monitor and control a process.
In business intelligence, a calculated value based on measurements used to monitor and control a process or business activity. Most metrics are ratios comparing one measurement to another.
See chart, metro map.
Software that allows applications to interact across hardware and network environments.
In project management, the marked end of a task or set of tasks, usually accompanied by some sort of event or a record of approval.
A measurement of processing speed. The concept is mistakenly considered a relative measure of computing capability among models and vendors. It is a meaningful measure only among versions of the same processors configured with identical peripherals and software.
A data store presenting a small subset of a data warehouse used by a small number of users. A minimart is a very focused slice of a larger data warehouse. See also data mart (DM).
The consumption of a high percentage of CPU cycles by a database query.
An exact copy of a dataset, kept up to date in real time. See also data replication.
Erroneous classification of a subject into a category in which the subject does not belong.
The value occurring most frequently in a range of values.
An abstract representation of how something is built (or is to be built), or how something works (or is observed as working).
A model of any kind that is independent of implementation and usage context, consisting solely of basic entities and relationships at a high level.
Generally, a very high-level block diagram listing the main terms and definitions for a business or system.
The storage and configuration management of models (including change control).
The process of adjusting a machine learning model to improve its accuracy, efficiency, or both.
A software development process that creates models or abstractions of a system or data in order to increase basic compatibility between systems.
An application design paradigm for object-oriented applications that separates the underlying “model” of business objects from the “view” of presentation interface objects and the “controller” events that users perform. Overlaying the controller functions on the view creates the illusion of direct manipulation.
In artificial intelligence, a class of computational algorithms that rely on repeated random sampling to estimate results, often used in reinforcement learning.
In general, a person’s internal rules of behavior. Contrast with ethics. See also professional ethics.
A problem-structuring and problem-solving technique that reduces the parameters to a finite number with a finite number of possible values, and then compares them to one another. Similar to construction of a dense cube with each parameter being a dimension. Also called Zwicky box after its inventor, Fritz Zwicky.
A computer language that allows one to specify which data to retrieve out of a multidimensional structure. The user process for this type of query is usually called slicing and dicing. The result of a multidimensional query is a cell, a two-dimensional slice, or a multidimensional sub-cube.
Systems with multiple interacting artificial intelligence agents that may cooperate or compete to achieve goals.
In physics and mathematics, describes an item that has a greater-than-two minimum number of coordinates necessary to specify it.
In data analysis, describes a data attribute that must be described by two or more distinct parameters.
A group of data cells arranged by the dimensions of the data. For example, a spreadsheet exemplifies a two-dimensional array with the data cells arranged in rows and columns, each being a dimension. A three-dimensional array can be visualized as a cube with each dimension forming a side of the cube, including any slice parallel with that side. Higher-dimensional arrays have no physical metaphor, but they organize the data in the way users think of their enterprise. Typical enterprise dimensions include time, measures, products, services, and geographical regions.
A query and calculation language designed for online analytical processing cubes, enabling complex analysis across multiple dimensions such as time, geography, and product hierarchies. From Multidimensional Expressions.
A data structure consisting of multiple files that may also include the explicit definition of relationships between/among the files. See also data model, physical (PDM).
Storage devices for multimedia files that also contain applications to display or play the multimedia files.
An artificial intelligence technique that integrates and processes information from multiple modalities (e.g., text, image, audio).
Characteristic of a relationship as either at most one (exclusivity) or more than one.
A model showing evaluation based on multiple variables.
To transform data such that the original data is unrecognizable without knowing the transformation rules and sequence, which are unpredictable or inconsistent. Sometimes accomplished with substitution of characters in order to obfuscate the original data. Occasionally explained as “modify until not guessed easily.”
A modeling technique in which each attribute describes a class independent of any other attributes that also describe that class. Named for Thomas Bayes. See also predictive modeling.
Generally, the designation of an object by a linguistic expression.
In data modeling, a class word, abbreviated usually to nm.
A pattern of assigning names, words, or parts of words to objects, often intended to convey meta-information that promotes consistency and ease of use while avoiding conflicts.
Relating to n (some number) of entities in a relationship, the number of attributes or columns in an entity table, the number of arguments or operands that a function requires, or more specifically, the number of objects in a predicate in object-relational management.
A nonregulatory agency of the US Department of Commerce. The institute’s mission is to promote US innovation and industrial competitiveness by advancing measurement science, standards, and technology in ways that enhance economic security and improve quality of life. As part of this mission, NIST awards the Malcolm Baldrige National Quality Award. Formerly known as the National Bureau of Standards.
A combination of business-oriented columns in a table that provides a candidate primary key.
See also key, business.
The process of describing information using proper sentences, not abbreviations or sentence fragments.
A field of artificial intelligence that focuses on the interaction between computers and humans through natural language. The goal is to enable computers to understand, interpret, and generate human language.
Transmission with a minimal amount of propagation and buffering delays.
Data that is not on line but is capable of being accessed and placed on line within 15 seconds of the access request. Archived data may be kept in nearline storage. See also archive.
An attribute of a relation, itself representing a relation. In a relational database management system, a column that contains a table in each row.
A comparison of the current value of a dollar versus the value of a dollar at some future time, after allowing for future influences such as inflation and expected rates of return. Positive NPV is said to indicate a good investment.
Visually, a graph of nodes and connections where more than one entry point for each node is allowed.
In data architecture, a topological arrangement of hardware and connections to allow communication between nodes and access to shared data and software.
The American National Standards Institute (ANSI) standard (first adopted in 1986) based on the CODASYL Network data structure. It was substantially rolled into the SQL:1999 ANSI standard, which is no longer relational. CODASYL stands for Conference on Data Systems Languages. In data management, it is best known for defining the CODASYL database model, also called the network data model.
An addressable device or connection point attached to a network.
A form of storage that attaches to a network but does not provide server functions such as file management.
A series of algorithms that attempt to recognize underlying relationships in a set of data through a process that mimics the way the human brain operates.
A form of artificial intelligence that uses algorithms to generate artificial neural networks optimized for a particular task.
A fundamental component of an artificial neural network, usually representing an input, a meta-feature, or an output. Neurons in a network-based artificial intelligence system are organized in layers.
A marketing segmentation strategy in which the firm focuses on serving one segment of the market. Similar to segmented marketing, but a niche is a small distinguishable segment that can be uniquely served.
A data modeling technique; the predecessor to object–role modeling. Named for developer G. M. Nijssen. Also called Natural Information Analysis Method.
In graph theory, a generic representation of something in a graph; it could be a type (representing a population) or an individual instance. Usually represented by some icon (e.g., box, circle) in the diagram.
In the context of neural networks, a basic processing unit (also called a neuron) that receives input, applies a transformation (activation function), and passes the result to the next layer.
Unwanted sound or data included with or around wanted sound or data
Random or irrelevant data in the input or output that can reduce the performance of artificial intelligence models if not properly handled.
A systematic naming of things or a system of names or terms for things.
In classification, a systemic naming of categories or items.
A number system that has no arithmetic or ordering significance and hence can only be compared as match or no match. Other operators are meaningless: multiply, divide, add, subtract, comparative (<, =, . . .), or Boolean. Probably the most commonly occurring type of numerical data in databases. Examples include account numbers. Often used as codes for particular characteristics or values in the real world.
An agreement between parties to not share specific confidential information without proper authorization from other involved parties.
A set of data in context that is not relevant or timely to the recipient.
A mathematical distribution of points around an axis that represents the mean of the dataset values and resembles a bell (low at both ends and high in the middle).
A characteristic of a file or table that indicates that it satisfies one or more of the rules of normalization. Not all rules must be satisfied in order. See also normalize.
A level of normalization where exist no multi-valued dependencies within a record or row; all attributes are atomic (single valued) for each entity instance. Multi-valued attributes and repeating groups must be removed from the record. In practice, any relational database table with a primary key assumes first normal form.
Describes a table structure that satisfies the six properties of a relation:
- All rows are unique.
- Order of rows is unimportant.
- All columns have unique names.
- Order of columns is unimportant.
- All values in a column are the same type.
- No column contains multiple values in the same row.
A level of normalization where every non-key attribute is fully dependent on the key in its entirety. In practice, when entities have compound keys, seek out any attribute that is dependent upon only part of the key and create a separate entity for what is identifiable by anything less than the whole key. Violations are commonly known as partial key dependencies.
A form of normalization where every entity has no transitive dependencies, that is, every data item must be directly determined by the identifier and not indirectly determined through some other non-key attribute. For example, if the boss of a department was determined by the department, then it would be incorrect to store both the department ID and the boss name in the employee record. Violations are commonly known as inter-attribute dependencies. Sometimes colloquially referred to as “the key, the whole key, and nothing but the key.”
An advanced level of normalization where no instance contains two or more independent multi-valued facts. Violations are commonly known as derived data.
An advanced level of normalization where all attributes of a concatenated key are independent of each other and cannot be derived from the remainder of the key. Violations are commonly known as inter-entity dependencies.
An advanced level of normalization that adds temporal constraints upon relations, such as a range of time when a relationship was effective. See also database, temporal.
Sometimes used incorrectly to refer to domain-key normal form; not the same as definition 1.
A level of normalization where every attribute or combination of attributes that can uniquely identify an instance is identified as a candidate key. Violations are known as overlapping keys.
A level of normalization that requires that a database only contains key constraints and domain constraints.
A level of normalization where there are no candidate keys that reuse the same attribute.
Generally, to impose standards or regulations, or to bring to a desired state.
In data modeling, to apply rules to a record-based data structure to reduce redundancy, such that each data attribute is stored
- as few times as necessary, and
- with its determinant as the identifier.
The rules of normalization are applied only within a record or table and cannot be applied until an identifier is first designated for the table. Even though the rules of normalization are numbered, there is no necessary ordering. They can be applied in any order, and some may be satisfied while others are not. For example, a record may have no transitive dependencies (thus not violating the condition for 3NF) but may have a partial dependency (thus failing 2NF).
A model that describes how a system should work according to assumptions or predefined standards.
A taxonomy of business classification; the standard used by federal statistical agencies in classifying business establishments for the purpose of collecting, analyzing, and publishing statistical data related to the US business economy. Replaces the Standard Industry Code. NAICS was developed jointly by the US Economic Classification Policy Committee, Statistics Canada, and Mexico’s Instituto Nacional de Estadística y Geografía to allow for a high level of comparability in business statistics among the North American countries.
A type of database that is distributed to enable large-scale data access. NoSQL does not use the traditional tabular relationship model, making it more flexible than relational databases.
A type of word that describes a person, a place, a thing, or an idea. One of the syntactic components used to construct sentences according to a grammar.
The absence of any value. A null value conveys that the value does not exist. It does not denote why the value is missing. Placing a zero or blank in the row would not reflect the accurate state of the row because zero and blank are values.
Generally, the prediction that an observed result is not due to any inherent systemic cause.
In data analysis, the prediction that one variable has no association with and responds independently of another variable.
Generally, to conceal through confusion.
In data security, the process of permanently scrambling or replacing data with unrelated values in order to conceal the original data permanently. Used to remove sensitive information from data when being transferred to unsecure systems.
In the real world, a person, place, thing, or concept. See also entity; instance.
In an object-oriented design, an instance of a class or a population of objects or events.
In an object-oriented program relating to object type, the code in memory that describes the attributes and allowable behavior of a business objects, interface object, or control object.
A discrete entity that encapsulates both data and behavior, commonly used in object-oriented data modeling and system design.
Generally, a set of ideas, abstractions, or things in the real world that can be identified with explicit boundaries and meaning, and whose properties and behavior follow the same rules.
Specifically, the definition of a set of objects that conforms to that definition.
In an object-oriented design, a collection of objects (instances) that conform to the same definition of structure and behaviors.
A data model that represents information as objects, including their attributes, relationships, and behaviors.
A computer vision task that identifies and locates objects within an image or video.
A unique value assigned to an object in order to track it simply and efficiently. Generally system assigned, immutable, and not visible to the user/programmer (unlike keys in a relational or entity-relationship database). Used to establish and maintain the integrity of defined relationships within the database.
A nonprofit organization that promotes object-oriented technology and open systems standards.
A collection of objects or classes.
The description of an object’s properties.
Generally, a form of design organized around objects (instances) where objects can be built (re)using other similar objects. For efficiency, the notion of object class was added to define a set of objects only once. See also object.
In data management, a style of software development (analysis, design, programming, and testing) organized around classes of objects in which the code encapsulates the data. object-oriented approaches promote data hiding, cohesion, class inheritance, and reuse.
A development environment for designing, building, and testing software using object-oriented languages and techniques.
A data model built up from elementary fact sentences; could also be thought of as a “no file” modeling scheme because it is not record-based. The main modeling constructs are objects and relationships. Objects encompass both entities and attributes. An object has attributes by virtue of the role that it plays in relationship with other objects. See also table think; Nijssen’s Information Analysis Method (NIAM).
A data storage architecture that manages data as objects, typically including metadata and unique identifiers for each object.
A specific, quantified target of achievement against which progress toward attainment can be measured. Achieving an objective contributes to achievement of a more general goal. A good objective is SMART: simple, measurable, attainable, realistic, and timely.
The practice of not including personal biases or preferences during an evaluation; evaluating on the agreed-to standards and facts alone.
A recorded data value or measurement collected during data acquisition, analysis, or monitoring activities.
Generally, an event; the fact that an event happened.
In data management, a physical record, row, or document representing an entity instance.
In the data resource, a set of entities in mathematics.
A specific record selected from a set of redundant records as the authoritative record, into which data from the other records can be consolidated.
A numbering system using a base of 8.
Data storage on a physical medium that is disconnected from any network or computer system.
Storage of digital information on devices or media that are inaccessible via online connections; often used for security, backup, and long-term storage. The data can only be accessed when connected to a computer or other device.
An information asset formally recognized by an organization as evidence of business activity and subject to records management requirements.
The characteristic of a relationship in which a member of population A must be related to only one member of population B, and vice versa. See also cardinality; relationship.
The characteristic of a relationship in which a member of population A must be related to one or more members of population B, but not vice versa. See also cardinality; relationship.
The characteristic of a relationship in which a member of population A may be related to one or more members of population B, but not vice versa. See also cardinality; relationship.
The collection of structures and processing routines to store and manipulate a dimensional data model (whether stored as a cube or a star) to enable multidimensional analysis of business trends and development of business projections. The term was originally coined by E. F. Codd. From online analytical processing. Opposite of online transaction processing (OLTP).
A class of systems designed to support analytical questions and multidimensional data analysis.
An end-user application that can request slices from online analytical processing servers and provide two-dimensional or multidimensional displays, user modifications, selections, ranking, calculations, etc., for visualization and navigation purposes. May be as simple as a spreadsheet program retrieving a slice for further work by a spreadsheet-literate user or as high-functioned as a financial modeling or sales analysis application.
Online analytical processing in which the data to be analyzed is stored on a desktop computer rather than on a conventional storage system. Also called desktop OLAP.
Online analytical processing that can provide multidimensional analysis simultaneously of data stored in a multidimensional database management system and in a relational database management system. Also called hybrid OLAP.
A Java application programming interface (API) for the Java 2 Platform, Enterprise Edition environment that supports the creation, storage, access, and management of data in an online analytical processing application. Hyperion, IBM, and Oracle initiated the development of JOLAP, intending it to be a counterpart to Java Database Connectivity specifically for OLAP. Also called Java OLAP.
Online analytical processing that only uses a multidimensional database management system to drive analysis. Also called multidimensional OLAP.
A version of online analytical processing where data is stored in random-access memory (RAM) rather than on disk, and calculations are performed on-the-fly, rather than stored. RTOLAP has a limitation of size because all data must be stored in RAM and space is at a premium; calculation results are therefore not stored. Also called real-time OLAP.
Online analytical processing that performs multidimensional analysis on data stored in a relational database management system. The multidimensional processing may be done within the RDBMS, a mid-tier server, or the client. A “merchant” ROLAP is one from an independent vendor that can work with any standard RDBMS. Also called relational OLAP.
Online analytical processing where spatial data is included in the data to be analyzed. Also called spatial OLAP.
Online analytical processing where the interaction with the data is through a web browser and may include the addition of web-based applications when analyzing or displaying the data. Also called web OLAP.
A machine learning approach where the model learns incrementally from data as it arrives, rather than all at once.
Data storage that is directly accessible to systems and users for immediate processing or retrieval.
A class of systems optimized for managing high-volume transactional data with frequent inserts, updates, and deletes. From online transaction processing.
Generally, the grammar rules for usage of a controlled vocabulary to create meaningful expressions within a domain or subject area.
In data management, a semantic data model defining structure and meaning, typically used to model non-tabular data. See also schema.
A formal representation of knowledge as a set of concepts and the relationships between them, used in artificial intelligence for reasoning and data integration.
In semantic modeling, a standard for defining ontologies.
A philosophy and practice requiring that some data be freely available to everyone, without restrictions from copyright, patents, or other mechanisms of control.
A publicly available specification that defines how data should be structured and exchanged to support interoperability.
A standard to allow programmers to write to an abstract relational database layer and delay binding until run time. Developed by the SQL Access Group consortium, ODBC has been widely adopted with modifications by Microsoft.
Free learning materials available via the internet.
An international voluntary standards group with more than 300 commercial, governmental, nonprofit, and research member organizations worldwide, collaborating to develop and implement standards for geospatial data and services, and geographic information system (GIS) data processing and exchange. Previously known as Open GIS Consortium. OGC specifications include the Web Map Service (WMS), Simple Features SQL (SFS), and Geography Markup Language (GML).
A detailed method and set of supporting tools for developing an enterprise architecture, developed by the Open Group.
An organization that promotes the adoption of standards for UNIX operating systems.
In software, freely available code, meaning the customer can download, install, begin using, or customize the code without paying.
Artificial intelligence tools and frameworks whose code is publicly available for use, modification, and distribution.
A form of data acquisition and analysis that focuses on publicly available data.
Course materials from learning institutions made available on the internet for free.
A high-level description of how an organization structures roles, responsibilities, processes, and technology to deliver data management capabilities.
A guiding statement that informs decisions and behaviors related to data management practices.
In the DAMA-DMBOK Functional Framework, a service and support activity performed on an ongoing basis.
The application of business intelligence (BI) tools to provide BI to the front lines of the business, where analytical capabilities guide operational decisions. Used to manage and optimize business operations.
An integrated database of operational data. Its sources include legacy databases and other operational databases. Contains current or near-term data; may contain 30 to 60 days of information, while a data warehouse typically contains years of data. As in a data warehouse, data in an ODS is extracted from sources, cleansed, consolidated, and transformed into a standard format. An ODS supports enterprise reporting, master data management, and application integration as the enterprise source for shared operational data. May serve as the primary source for a data warehouse or be used to audit a data warehouse.
An operational data store where data is moved from sources almost immediately after being written in the source system, without any integration or transformation.
An operational data store where data is moved from sources within a few hours after being written in the source system, allowing for some integration and transformation.
An operational data store where data is moved from sources overnight, following being written in the source system, allowing for integration and transformation.
An operational data store where summary data is moved from a data warehouse into the ODS for operational use.
A defined point of interaction through which systems exchange operational data.
Metadata that describes the technical and operational aspects of data processing, including lineage, scheduling, and system usage.
Measurable outcomes relative to stated enterprise-wide operational goals.
The production of routine reports that support the monitoring and managing of ongoing business operations.
The risk of loss of failure resulting from inadequate or failed processes, systems, or data.
The practices and controls used to protect operational data and systems from unauthorized access or disruption.
An application that runs the business on a day-to-day basis using real-time data, typically online transaction processing (OLTP) databases.
The process of implementing data, analytics, or models into operational systems and business processes.
Technology that can scan typed or handwritten characters and convert them to digital text, usually ASCII.
The process of capturing data from document forms that have been human-generated rather than computer-generated, such as highlights and margin notes.
To configure a system to perform more in accordance with some expected measurement than another configuration.
An artificial intelligence system designed to perform optimalization.
The process of improving performance, efficiency, or quality by adjusting processes, systems, or models.
An artificial intelligence process aimed at finding the best value for a function, given a set of restrictions. Optimization is a key in all data science systems.
Generally, not required. Opposite of mandatory.
Characteristic of an attribute where a value is not required by an entity constraint (NULLs allowed in SQL).
Characteristic of a relationship in which an entity or object instance need not relate to any member of the other entity type population (i.e., can be an orphan).
In data modeling, the minimum number of times one entity can be associated to another entity. The choices are either zero or one.
Coordinating automated tasks and data flows across systems and platforms.
Generally, the sequence of items or events in time or ranked by some quality, such as importance.
In data services, a message sent that triggers the delivery of required data. There are three types of orders: select order, transform order, and propagate order.
A number that signifies sequence within a set, or a rank, solely for comparison or matching. Does not signify quantity and cannot be meaningfully added or subtracted.
In general, an arrangement of people, dedicated to common goals, who control the organization’s performance and have a clear delineation of what is included in the organization.
One of the DAMA-DMBOK Functional Framework Environmental Elements. Includes management; critical success factors; reporting structures; contracting strategies; budgeting and related resource allocation issues; teamwork and group dynamics; authority and empowerment; shared values and beliefs; expectations and attitudes; personal style and preference differences; cultural rites, rituals, and symbols; organizational heritage; and change management recommendations.
A not-for-profit consortium that advances e-business by promoting open, collaborative development of interoperability specifications.
Policies, procedures, and structures used to direct and manage data-related activities within an organization.
A coordinated plan that defines how data will be managed, governed, and used to support organizational objectives.
The collected data of the enterprise about itself and its environment, in current context.
Information that is of significance to the organization, is combined with experience and understanding, and is retained. Information in context with respect to understanding what is relevant and significant to a business issue or business topic; what is meaningful to the business.
A model showing the organization of a particular system or company.
Literally, to be at right angles. Typically refers to characteristics that are as independent of each other as possible. For example, data and processes are considered orthogonal to each other.
A data instance that is extremely deviated from the mean of the rest of the dataset.
Identifying data points that differ significantly from the majority, often used in anomaly detection.
A dimension that implements denormalization by qualifying another dimension. This results in a star schema.
The process of arranging services to be done by an external party, to replace the need for an internal party to perform those services.
A markup language for showing terms and relationships in vocabularies.
In software, a predeveloped application software product available for purchase.
In object-oriented software, a unit of deployment, usually consisting of many related object-oriented classes.
Value-added solutions with embedded knowledge of business processes and specific functional metrics based on industry best practices, available for purchase.
The process of splitting datasets into finite blocks (pages) for optimal storage performance.
The process of retrieving and/or swapping parts of datasets (pages) as they are required.
An example of pattern that represents an acquired way of thinking about something that consciously and/or unconsciously shapes thought and action.
The ability to perform multiple functions in parallel.
An attribute of many systems and algorithms, whereby different parts of them can be split among various central processing units or graphics processing units working in parallel. Often essential for artificial intelligence processes.
In data management, a data attribute provided as input to a system or process.
A single bit that represents the count of the preceding bits that equal 1 in value. Used to check data transmission: If the parity bit says there were an odd number of 1 values, and the data shows an even number of 1 values, then there is an error in transmission.
To analyze a sequence using predetermined rules to determine content or value.
In general, to split into parts according to some rule or condition.
To logically and/or physically segregate data in a single table into multiple files, each containing groups of similar rows that are more easily maintained or accessed. Relational database management systems typically provide this functionality. Partitioning of data aids in performance and utility processing.
One segment of a dataset identified by a specific condition.
A method of partitioning a table horizontally using one partitioning method first and then partitioning the resulting set using another partitioning method. Common types are range-list and range-hash.
An attribute or expression used to differentiate parts of datasets.
A method of partitioning a table horizontally in which the partitions are identified by a hash value derived from one or more columns in the table.
A method of partitioning that divides a single logical table into multiple physical tables based on the row values of the primary key column. All columns generally appear in each table, but each table contains a subset of the logical table’s rows (either discrete or overlapping subsets). Employed when there is a regular need to access or isolate a readily identifiable subset of the rows to meet security, distribution, and performance optimization needs. Note: It is horizontal only because of the convention used to represent a table, namely, columns across the top and rows down.
A method of partitioning a table horizontally in which the partitions are identified by presence of a column’s value in a list of possible values.
A method of partitioning a table horizontally in which the partitions are identified by the upper and lower bounds of one or more columns in the table.
A method of partitioning that segregates the columns of a single logical table into multiple physical tables. All logical rows may appear in each new table, but each new table contains a subset of the original table’s columns. Some columns may be redundant across tables and will necessarily be so for primary key columns. Vertical partitioning is employed when there is a regular need to access or isolate a readily identifiable subset of the parent table’s columns. This technique may be effective to meet security, distribution, and usability requirements. Note: It is vertical only because of the convention used to represent a table, namely, attributes across the top and entity instances down. See also table, outrigger.
A string of characters used to help authenticate a user logging into a system.
A series of one or more arcs between nodes in a graph.
The process of identifying patterns and regularities in data.
A worldwide information security standard assembled by the Payment Card Industry Security Standards Council.
Measurable outcomes relative to stated goals.
Responsibility assumed for achieving objectives and disclosing present and future variances against those objectives.
Notification via email, portal, or wireless device of a key trend or business event that is associated with an objective.
Activities related to understanding and improving computer hardware and software performance (response time and throughput), including database performance.
A strategic management process designed to translate an organization’s mission statement and overall business strategy into specific, quantifiable objectives and to monitor the organization’s performance in terms of achieving those objectives.
Generally, the interval of single repetition of a varying quantity, motion, or phenomenon that repeats itself regularly.
Specifically, a quantity of time.
The frequency of compilation of the data (e.g., a time series could be available at annual frequency, but the underlying data is compiled monthly, thus setting a monthly periodicity).
A state or status that lasts beyond the process that created it.
Data that outlasts the execution of a particular program, stored in the records of the enterprise and available for reuse.
Permanently and irreversibly altering data by masking. Typically used in development rather than production.
Applies to every organization that collects, uses, and disseminates personal information in the course of commercial activities.
Information that refers to a specific individual. Includes name, address, telephone number, governmental identification numbers, US Social Security numbers, etc.
A ubiquitous, wireless, always-on, networked world.
A way hackers redirect users to false websites by means of a virus that changes or modifies the user’s host files.
A form of social engineering in which a hacker sends emails that appear to come from legitimate sources in order to deceive the recipient into revealing personal or sensitive information or installing malware such as viruses, worms, adware, or ransomware.
The act of developing a physical data model.
A series of data processing steps or stages, often used in extract-transform-load or extract-load-transform.
To rotate the view of data. Used in multidimensional analysis with online analytical processing tools, but can also be performed in spreadsheet applications.
A multidimensional modeling scheme (specifically found in Microsoft Excel and many business intelligence tools).
In general, to define goals and objectives and to devise approaches and activities to realize or achieve these goals.
In information services, to define mission and purpose statements, goals, objectives, critical success factors, strategy, architecture, programs, and projects for an enterprise, and then to assess and analyze to guide decisions. Often considered the first phase in the software development lifecycle, although occurring before project initiation.
An organized set of goals, objectives, and activities.
A circular process for continuous improvement. Also called the Shewhart cycle, after its developer, Walter A. Shewhart. See also Deming cycle.
A sequence of steps to retrieve or manipulate data that outlines the order in which tables are accessed, filters are applied, or joins are made.
In the DAMA-DMBOK Functional Framework, an activity that sets the strategic and tactical course for other data management activities. Planning activities may be performed on a recurring basis.
Any base of technologies on which other technologies or processes are built and operated to provide interoperability, simplify implementation, streamline deployment, and promote maintenance of solutions. The platform resource consists of hardware and system software.
A software package delivered as a service that allows third-party applications to “plug in.” Facebook and X (formerly Twitter) are examples.
A data type that serves only to refer to another data point’s storage address.
A distribution curve where the tail on one side is longer and thinner than the other. Named for Siméon Denis Poisson.
A statement of a selected course of action and high-level description of desired behavior to achieve a set of goals.
In object-oriented design, the implementation of subclasses of a parent class so that identical requests sent to different child classes are handled differently without the caller knowing.
A collection of things (instances) that are considered part of the same set, called a type.
In general, a collection of things (instances) which are considered part of the same set, called a type.
The process of loading and replicating multiple rows of data into a relational database on a one-time or recurring basis. See also data loading; data replication.
A website designed to be the “front door” through which a user accesses links to relevant sites. Typically, a portal site has a catalog of sites, a search engine, or both. A portal site may also offer email and other services to entice people to use that site as the main point of entry or portal to the web.
A collection of assets, liabilities, and/or issues to manage.
A notation where position affects the value of a character or digit. Binary, octal, decimal, and hexadecimal are all examples of positional notation. Also called place-value notation.
A repeatedly performed, customary way of doing something.
One of the DAMA-DMBOK Functional Framework Environmental Elements. Common and popular methods and procedures used to perform the processes and product the deliverables. Practices and Techniques may also include common conventions, best practice recommendations, and alternative approaches without elaboration.
The level of detail of a data attribute, usually expressed as the number of numeric places to the right of a decimal point. See also scale.
A performance metric in classification tasks: the proportion of true positive results among all positive predictions.
Generally, a statement that can be evaluated as true or false. For example, WHERE clauses of SQL SELECT statements define predicate logic for qualifying rows. See also arity.
In object–role models, a labeled relationship on one or more objects. Depending on the number of objects, a predicate may be unary, binary, ternary, etc.
The estimation of future results or other dataset results based on existing data.
The output or result that an artificial intelligence model generates based on input data.
Methods of directed and undirected knowledge discovery, relying on statistical algorithms, neural networks, and optimization research to predict and recommend actions based on discovering, verifying, and applying patterns in data to predict the behavior of customers, products, services, market dynamics, and other critical business activity.
An area of statistical analysis that deals with extracting information from data and using it to predict future trends and behavior patterns.
The discipline of getting to know customers (or citizens) by performing complex analysis (including data mining) on customer data.
The process of estimating the probability of a specified outcome given an input dataset.
The step of cleaning and preparing data before feeding it into an artificial intelligence model.
An encryption program.
One of the DAMA-DMBOK Functional Framework Environmental Elements. The information, physical databases, and documents created as interim and final outputs of each function. Some deliverables are essential, some are generally recommended, and others are optional depending on circumstances.
A word used in the name of an attribute to identify its domain (logical datatype). See also class word.
In general, simple, unsophisticated, and/or uncomplicated.
In data modeling, an entity or class that has no supertypes. There is disagreement over whether there are just a few semantic primitives of which all other entities can be considered subtypes.
Formally, a fundamental law, doctrine, premise, or assumption.
Informally, a rule or code of conduct.
Free the database of modification anomalies.
Minimize redesign when extending the database structure.
Make the data model more informative to users.
Avoid bias toward any particular pattern of querying.
In data security, the need for access control and usage monitoring.
Protection of personal or sensitive data from unauthorized access.
Unavailable for observation at all, or only to a limited set of observers. See also confidentiality.
Opposite of public.
A type of matching that relies on statistical analysis of a sample dataset to project results on the full dataset.
A model that incorporates uncertainty and provides outputs in the form of probability distributions.
Generally, a series of low-level steps or tasks in a process followed in a defined and repeatable order.
In data management, a set of instructions for human users of computer systems that augments the automated workflow.
Generally, an action (or set of related actions in a value chain) occurring to accomplish something. Functions, activities, procedures, steps, and tasks are subtypes of process. The execution or carrying out of a process constitutes behavior. Not the same as a functionally similar grouping of actions; the actions have to have a logical progression or relationship.
The systematic evaluation of the performance of a process, taking corrective action if performance is not acceptable.
Specifies methods for business and systems planning, analysis, and design processes.
The analysis, control, and improvement of a business process and its interrelated steps.
The person responsible for process definition, execution, and control.
The definition or specification of how a process is to be carried out. A computer program is a process specification, to be carried out by the computer (the processor).
Generally, something produced. The output or result of a process. Something tangible, as opposed to a service. Synonymous with an output, result, or deliverable.
Solutions for capturing and maintaining accurate, up-to-date data about an organization’s products and delivering information in an actionable form “just in time” at product development or distribution points. A specialized form of master data management focusing on product master data.
Processes and tools used to predict and evaluate success of products through marketing and sales efforts.
An occupational calling (vocation) requiring specialized knowledge.
The body of persons engaged in that vocation.
A designation earned by a person verifying that the individual has the knowledge, skills, or abilities that qualify them to perform a job. While licensing is required by law, certification is generally voluntary. Professional certifications are awarded by certification body, usually a professional organization. People become certified through training and/or passing an exam. Individuals often advertise their status by appending the abbreviation for the designation to their name. See also profession.
Training, mentoring, and continuing education in a professional field of study to attain, maintain, and extend one’s mastery of professional skills. See also profession.
Principles of standards of conduct with which all members of that profession are expected to comply. See also ethics; morals.
Analyzing data for quality, structure, and content, often to identify anomalies. Also known as data profiling.
A set of projects that addresses a common set of goals and objectives; a long-term initiative made up of several parallel or incremental projects.
A model for project or process management to evaluate tasks involved in the project or process in order to find the shortest duration possible. Also called pert chart.
The planning, supervision, and control of a program.
An effort with a defined purpose, start, and finish.
The planning, supervision, and control of a project.
A nonprofit organization of project management professionals. PMI is the sponsor of the PMBOK® Guide and the certifying body for the Project Management Professional certification.
A detailed description of a proposed effort.
A minimal implementation or execution of a process that serves as a sample sufficient to prove the success of the whole implementation or process.
Data that is transferred from a data source to one or more target environments according to propagation rules normally based on transaction logic. See also data replication.
An attribute or a relationship of an object.
Data created or acquired by an organization that is not publicly available and provides a competitive or economic advantage to its owner. Examples are intellectual property, client lists, and financial data.
Any individually identifiable health information that is created, stored, transmitted, or received in any form and is protected under the United States HIPAA Privacy Act. This law gives individuals rights over their health information and sets rules and limits on who can look at and receive individuals’ health information. The Privacy Rule applies to all forms of individuals’ protected health information, whether electronic, written, or oral. The Security Rule, a US federal law that protects health information in electronic form, requires entities covered by the Health Insurance Portability and Accountability Act (HIPAA) to ensure that electronic protected health information is secure. Also called personal health information.
A set of conventions that governs the communications between processes. Specifies the format and content of messages to be exchanged.
An artifact in iterative development. A prototype may be disposable or the base for further incremental development.
To create a test artifact for the sole purpose of determining whether the design is feasible or will be successful, given environmental restraints.
Representing the origin or source of something, the history of ownership, the location of an object. The term is used in a wide range of fields, including science and computing.
In customer relationship management, a segment of a population delineated by certain shared preferences, activities, or attitudes.
The act of making information or data readily accessible and available to all interested individuals and institutions.
Works that have no copyright restrictions on them, are freely available, and are usable without restriction.
Encryption technologies and services designed to protect the security of communications and business transactions on the internet.
The entity or organization that makes something available for common use.
A DCMI element in element set Intellectual Property: an entity that provides accessibility to a resource. See also Dublin Core Metadata Initiative (DCMI).
Generally, to remove, cleanse, or empty.
In data management, to permanently delete data. See also archive.
The types of movement of things or data between two systems or entities. The system or entity that may push; the system or entity that consumes may pull.
A high-level, general-purpose programming language employed in a wide range of applications.
A composite metric representing overall data quality across multiple dimensions, such as accuracy, completeness, and timeliness.
An assessment ensuring systems, data, or processes meet predefined qualification standards.
A rule determining whether data is acceptable for use in a specific context or analysis.
Data that has passed defined validation, quality, and compliance checks.
Analysis focused on patterns, meanings, and themes rather than numeric measurement.
Non-numeric data representing characteristics, categories, or attributes.
Systematic activities implemented to ensure that processes meet defined quality standards.
A documented approach outlining quality standards, procedures, and responsibilities.
A characteristic used to evaluate data quality (e.g., accuracy, completeness, consistency).
A record documenting quality-related changes, decisions, and activities.
A reference standard used to evaluate data quality performance.
A proactive approach where quality is built into systems, processes, and data from inception.
A specific test, rule, or validation applied to data to ensure conformity to requirements.
Adherence to internal or external quality standards, regulations, and policies.
Operational techniques used to verify that outputs meet quality requirements.
A visualization used to monitor process stability and identify anomalies or trends.
Documented steps for inspecting, validating, and verifying data quality.
Defined conditions that data must satisfy to be considered fit for purpose.
A visualization platform presenting data quality key performance indicators and metrics.
A failure of data to meet one or more quality requirements.
A specific aspect of data quality, such as accuracy, completeness, consistency, or timeliness.
A recorded instance where data does not meet defined quality thresholds.
A structured model defining quality dimensions, metrics, roles, and governance processes.
A checkpoint in a data pipeline where quality requirements must be met before progression.
Oversight structures, roles, and processes ensuring accountability for data quality.
An iterative process for identifying, addressing, and preventing data quality issues.
An event where a data quality failure impacts operations, analytics, or decision-making.
A repository documenting data quality problems, root causes, and resolutions.
A key performance indicator (KPI) used to track data quality objectives.
Coordinated activities to direct and control an organization with regard to data quality.
A documented approach outlining quality objectives, standards, controls, and responsibilities.
A formal system documenting processes, responsibilities, and standards related to quality.
Metadata describing data quality characteristics, metrics, thresholds, and scores.
Continuous observation and tracking of data quality metrics over time.
A specific, measurable goal related to improving or maintaining data quality.
A formal statement outlining organizational commitment to data quality principles.
The process of examining data to understand quality characteristics, distributions, and anomalies.
Actions taken to correct data quality issues and prevent recurrence.
A summary of data quality findings, metrics, trends, and issues.
A defined expectation that data must meet to be usable for a given purpose.
A formal condition or constraint used to validate data values, relationships, or structures.
A consolidated view of multiple quality metrics for stakeholders.
A service-level agreement (SLA) defining quality expectations and performance measures.
A documented specification defining acceptable data quality levels.
A minimum or maximum acceptable value for a specific quality metric.
Verification that data meets defined quality rules, constraints, and expectations.
A statistical measure dividing a dataset into equalsized subsets.
Analysis using statistical, mathematical, or computational techniques.
Numeric data suitable for measurement and quantitative analysis.
Mapping continuous values into discrete levels.
A set of attributes that can indirectly identify an individual when combined with external data.
An analytical design lacking random assignment but used to infer causal effects.
A set of attributes that uniquely identifies records in most, but not all, cases.
Data that changes infrequently but is not fully static.
Data with partial structure, such as JSON or XML, not governed by rigid schemas.
A software layer that hides underlying query complexity from users.
A programmatic application programming interface (API) for executing queries against a data source.
A tool or interface used to construct queries without manual coding.
A visual query technique allowing users to specify queries using example data.
A measure of the computational and logical effort required to execute a query.
Estimated resource usage (e.g., CPU, memory, I/O) required to execute a query.
The strategy selected by a database engine to execute a query efficiently.
Executing a single logical query across multiple heterogeneous data sources.
A directive that influences how a database optimizer executes a query.
A user-facing system for submitting, managing, and reviewing queries.
A formal language used to retrieve, manipulate, or define data (e.g., SQL (Structured Query Language)).
The time taken for a query to return results.
Capturing executed queries for auditing, monitoring, or tuning purposes.
Software that routes, modifies, or manages queries between clients and data sources.
Techniques used to improve query performance and efficiency.
The route taken by a query through systems, services, or components.
Adjusting queries, indexes, or schemas to improve execution speed.
Stored execution plans reused to optimize query performance.
A condition used to filter data within a query.
The output produced by executing a query.
Directing queries to appropriate systems, nodes, or replicas.
The total duration of query execution.
A component that manages timing, prioritization, and concurrency of queries.
Rules controlling what data can be accessed or manipulated through queries.
Limiting query execution to protect system performance.
The collection of queries executed against a system over time.
Metadata that can be searched, filtered, and queried programmatically.
A data structure or system component managing ordered processing of tasks or messages.
The number of items waiting in a queue.
The delay experienced by items while waiting in a queue.
Administration and control of message or task queues.
A system state in which changes are temporarily halted to ensure consistency.
The minimum number of nodes required to agree for a distributed operation to succeed.
A defined limit on resource usage, such as storage, compute, or query execution.
Mechanisms ensuring defined resource usage limits are respected.
Technology for tracking the location of goods. RFID tags are transponders, devices that upon receiving a radio signal transmit one of their own. Transponders were first used during World War II as a means of identifying friendly aircraft, but now RFID technology has become economical for widespread uses, such as monitoring inventories. While their main function remains identification, they can also be used for detecting and locating objects as well as monitoring an object’s condition and environment.
A form of data storage that uses integrated circuits that allow changes in the data contents. Originally called read-alterable memory, to distinguish from read-only memory (ROM), which is still used.
Memory in which the time to access any one unit of information is the same as the time to access any other unit of information. Also called uniform access memory.
A restricted set of attribute values defined by a pair of minimum and maximum, or start and end, values.
Methods, tools, and techniques that dramatically accelerate application development time.
A process of reviewing environmental, resource, and work effort measurements in order to predict success of an implementation.
A periodic survey to determine user satisfaction with the environment and find out what information needs are not being met.
Permission for users to view data without being able to change it.
A form of data storage that uses integrated circuits that is written to once and then static afterward. Usually mass-produced in a factory as a secure distribution method for code and/or data.
A number that contains no imaginary components, of unspecified precision. Almost always displayed with decimal points. Contrast with integer.
The utmost level of timeliness regarding the use of information. Commonly defined as instant or instantaneous, although not strictly the same.
The condition in which the time to process and respond to a request for information (processing) is less than the time in which it is needed in the environment (the source of the request) to make a difference.
Up-to-the-second detailed data, used to run a business and accessed in read/write mode, usually through predefined transactions.
Expectations within specific operational contexts.
Generally, evidence of an organization’s activities. These activities can be events, transactions, contracts, correspondence, policies, decisions, procedures, operations, personnel files, and financial statements. Records can be physical documents, electronic files and messages, or database contents.
In data management, the physical representation of data about an instance. A collection of fields about an instance generally representing the information pertaining to an instance of a member of the type population.
The management of evidence of an organization’s activities. See also record, definition 1.
A list of types or categories of data that an organization holds, along with the length of time the records should be kept before being destroyed. Often outside regulations determine the length of record retention.
The ability to reestablish service after interruption and correct errors caused by unforeseen events or component failures.
Generally, the restoration of something to its status before an event or at a point in time.
In data management, the restoration of a database to its state as of a different point in time, typically in the wake of a hardware or software failure.
Restoration of a snapshot backup copy of the database (a valid snapshot copy of the data as of a point in time), followed by the re-execution of logged change activity since the backup copy was made. This method essentially reverses (rolls back) all changes after the snapshot was taken and re-executes from that point forward.
Restoration of a full backup copy of the database, followed by re-execution of logged change activity since the backup copy was made. This method essentially starts over from a full copy of the database and re-executes from that point forward.
The intent to recover data up to a specific point in a transaction stream following a downtime event. Expresses the amount of data loss an organization may tolerate.
The intent to recover lost applications, within specific time limitations, to ensure a certain level of operational continuity. Expresses the amount of time a business will tolerate the computing system (hardware, software, services) being offline.
A type of neural networks designed for sequential data (e.g., time series, natural language) where outputs from previous steps are fed back into the model.
Capable of being infinitely repeated, using one instance of the execution of a process as the input of the next instance of execution. A process calling itself.
Sometimes used to refer to a relationship in a data structure. See also relationship, reflexive.
Generally, exceeding necessary or normal measures; duplicative.
In data management, storage of multiple copies of logically identical data. Physically, the data may or may not be identical across systems, and it is not known which is most current or accurate.
Management of a distributed data environment to limit excessive copying, update, and transmission costs associated with multiple copies of the same data. Data replication is a strategy for redundancy control with the intention to improve performance. See also managed replication.
A technology for configuring a logical data storage device across multiple physical devices to improve performance, availability, or both. The primary goal is fault tolerance, as in most configurations data can be recovered after a device failure and, in some cases, without interruption.
An Agile term for improving, enhancing, or simplifying a previous coded program module.
It does not alter the external behavior of the code but improves its internal structure. Sometimes referred to as turning “dirty” code into “clean” code.
Ensuring consistency with a “golden version” of data values. Managing golden versions and replicas.
Generally, any data used to organize or categorize other data, or for relating data to information both within and beyond the boundaries of the enterprise. Usually consists of codes and descriptions or definitions.
In financial services, refers to both reference and master data together.
Processes that control vocabularies (defined domain values), including control over standardized terms, code values, and other unique identifiers; business definitions for each value; business relationships within and across domain value lists; and the consistent, shared use of accurate, timely, and relevant reference data values to classify and categorize data.
Reference data that includes geopolitical data, such as countries, states or provinces, counties, postal codes, sales territories, etc.
In data management, constraints that govern the relationship of an occurrence of one entity to one or more occurrences of another entity. These constraints may be automatically enforced by the database management system. For instance, every purchase order must have one and only one customer. If the relationship is represented using a foreign key, then the foreign key is said to reference a file or entity table where the identifier is from the same domain. Having referential integrity means that if a value exists in the foreign key of the referencing file, then it must exist as a valid identifier in the referenced file or table.
The condition that exists when all intended references from data in one column of a table to data in another column in the same or a different table are valid.
A process of taking a snapshot of data from one environment and moving it to another environment, overlaying old data with the new data each time.
Generally, a permanent collection of data related to some topic or collected through some process.
In metadata management, an application that stores metadata for querying and can be used by any other application in the network with sufficient access privileges.
Using one dataset to predict the results of a second.
A statistical method used in machine learning to model and analyze relationships between variables, especially for predicting continuous outputs.
A statistical technique that seeks a line which best fits through a set of data as plotted on a graph, finding the cleanest path that deviates the least from any instance within the set.
The testing of code that has been modified to ensure that previously developed and tested functions were not broken when the code was changed.
The act of meeting the requirements of government legislation or self-regulating industry organizational mandates. For instance, public companies are required to provide specific financial reporting and disclosure. Regulators in the United States include the securities authorities (the Securities and Exchange Commission [SEC]), tax authorities (the Internal Revenue Service [IRS]) and banking authorities (the Federal Deposit Insurance Corporation [FDIC]).
Generally, the manner in which two objects may be associated, ordered, connected, or otherwise grouped, using inherent attributes.
In data management, a physical structure (a flat file, inverted list, linked list, bitmap, hash table, b-tree, etc.) consisting of a set of one or more columns and zero or more rows. The relation is between the row category and column category.
A DCMI element in an element set: a relationship between any two resources or a resource and another instance of that same resource. See also Dublin Core Metadata Initiative (DCMI).
Generally, an instance of a connection between two or more things.
In data management, a link between two entities describing the business rules governing how the two entities interact in the real world, including their cardinality and dependency. Typically described using verb parts of speech.
In data modeling, a relationship between two entity types which itself has attributes. If the relationship is many-to-many, then the attributes on the relationship cannot logically be stored in either of the related entities.
In data modeling, a relationship that involves two entity types or object types. The relationship could be defined on a single entity type, in which case it is called a reflexive relationship. In such a relationship, the members play different roles in the relationship: for example, a boss–employee relationship in which all bosses are employees. Some have called a reflexive relationship unary because it involves a single population, but that is incorrect. It is still binary, with the members playing two different roles in the relationship.
In data modeling, a relationship where an instance of one entity is required, but an instance of the other entity is not required. Example: A product may not have any orders, but each order must have at least one product.
In data modeling, a one-to-many relationship between two entity types (which could be the same entity type; see relationship, reflexive), in which the entity type on the many side of the relationship is dependent upon the entity type on the one side; sometimes called a parent–child relationship. An instance of the child must relate to one and only one instance of the parent entity type.
In data modeling, a relationship where the child instance cannot be uniquely identified without knowing the parent instance or the identifier (key) of the parent instance in
that relationship.
In data modeling, a relationship where the both instances are required to be present. Example: An account must have an account holder. Each account holder must have at least one account.
In data modeling, a relationship where the child instance can be uniquely identified without knowing the parent instance or the identifier (key) of the parent instance in that relationship.
In data modeling, the particular graphical representation of a relationship and its characteristics. Most frequently, a relationship is represented by an arc drawn between two related things, with additional notations to reflect its characteristics. For example, the multiplicity characteristic of many (more than one) could be represented by a fork, an arrow, a double-headed arrow, an asterisk, or the letter M.
In data modeling, a relation instance where not all instances of either entity participate in the relationship. Example: A company location may not have assigned orders (a data center), and orders may not have assigned company locations (for a service done over the phone).
In data modeling, a relationship within processes in which a process calls itself during execution. Sometimes used to refer to a reflexive relationship in a data structure.
In data modeling, a relationship in a data structure in which individual instances are related to other instances of the same type (i.e., in the same file or table). For example, in an employee table, an employee could be related to some other employee who is their boss. Sometimes (erroneously) called a recursive relationship, which instead applies to a process calling itself.
In data analysis, a method for finding relationships between variables that exceed a rate of frequency or some other measure to determine a minimum significance. Also called association rule analysis.
In data modeling, a relationship that involves three entity types or object types. If only two of the participating entity types are required to uniquely identify instances of the relationship, then it would also be an attributed, binary, many-to-many relationship. For example, Employee, Skill, and Proficiency Level: For each Employee–Skill combination (a many-to-many relationship), if there can only be at most one Proficiency Level, then it can be viewed as a binary relationship between Employee and Skill, and Proficiency Level would be an attribute of the binary relationship.
A method of assigning pointer values based on a start point other than the root of the structure.
The process responsible for planning, scheduling, and controlling the movement of releases to test and live environments. The primary objective of release management is to ensure that the integrity of the live environment is protected and that the correct components are released. See also Information Technology Infrastructure Library (ITIL).
Generally, closeness of the initial estimated value to the subsequent estimated value.
In data management, the ability of a technology component (server, application, database, etc.) or group of components to consistently perform its functions within stated timeframes.
A technique used to create, distribute, and use Java objects.
A mechanism for invoking a service on another platform, or the communication using that mechanism that invokes execution of a subroutine or process on a different system.
A group of data items that together describe something; an attribute with multiple values within an instance of its parent entity. When related to some other entity in a “something-to-many” relationship, and stored in the related entity type, it becomes a sub-entity within that parent entity. Also called a nested relationship.
Generally, the process of making copies of something.
In data management, the copying of data from a data source to one or more target environments based on rules.
In data management, the state when data is replicated but users cannot see which of the duplicated systems is fulfilling the data request.
An automated business process or related functionality that provides a detailed, formal account of relevant or requested information.
Loosely used, any database or file. Not recommended for use. See metadata repository.
A managed metadata environment. See metadata repository.
A customer expectation of a product or service. May be formal or informal, stated or unstated, needed or desired.
A formal statement of need for data, functionality, or other characteristic.
Requirements stated in business terms or ordinary language what must be delivered or accomplished in order to return value.
A description of expected behavior of a system, given a defined set of inputs or events.
A description of expected operation of a system independent of any specific tasks or functions; may not be measurable in the same terms as other requirements. Includes reliability, efficiency, portability, etc.
The formal documentation of requirements, typically using standardized formats and templates, often stored in a requirements database for further analysis and validation/testing/verification.
The elicitation, specification, and modeling of requirements.
A term that has meaning outside of a computer language and therefore may not be used for other than its defined purpose.
The process of assigning system resources (e.g., memory, storage, compute) to data tasks or users.
The basic technique for expressing knowledge on the Semantic Web. Also called Resource Description Format.
Accountability for performance of a function, activity, or task by a role.
In object-oriented design, synonymous with a method.
A matrix used to describe classes of involved parties and their impact on a situation by role.
Generally, the process of keeping something in place.
In data management, the length of time that data is stored or archived before purging.
Rules that govern how long data is stored and when it should be deleted or archived.
The calculated financial return on a business initiative, comparing costs and benefits for a period of time.
The process of deriving a draft physical model representing an implemented system (application and/or database) from automated scanning of the implemented application and database objects, as a first step toward redesign.
Generally, the entitlements or freedoms that may or may not be acted upon by an entity.
In database management, the permissions to perform create-read-update-delete (CRUD) activities assigned to a user or role.
A DCMI element in element set Intellectual Property: rules regarding access to and through a resource. See also Dublin Core Metadata Initiative (DCMI).
A process to identify potential situations that could cause change to an effort from both internal and external forces; assign severity and priority ranks in order to determine overall risk; manage a situation or project to mitigate or minimize the occurrence of risk; and, if the risk materializes, to minimize loss or damage.
Managing a situation or project so that minimum loss or damage will result if a risk assessment materializes.
Defines the actions required to move from current to future (target) state. Similar to a high-level project plan.
The field of engineering and artificial intelligence focused on creating machines (robots) that can perform tasks autonomously or semiautonomously.
The ability of an artificial intelligence model to maintain performance when faced with noisy, incomplete, or adversarial data.
Generally, a label assigned to a set of connected behaviors, rights and obligations.
In data modeling, the way in which entities of one type relate to entities of another type in a relationship. See also object–role model (ORM).
A name used to refer to the logical set of related responsibilities assignable to a person or organization, and to parties with these assigned responsibilities.
The business and information technology roles involved in performing and supervising the function and the specific responsibilities of each role in that function. Many roles will participate in multiple functions.
To undo database statements performed prior to a commit of the transaction.
A query that summarizes data at a level higher than the previous level of detail.
A forecasting method that shifts planning away from historic budgeting and forecasting and moves it toward a continuous predictive modeling method. It requires access to relevant information from multiple data sources as well as business processes throughout the enterprise. Rolling forecasts can be updated continuously throughout the year to improve accountability.
The underlying fundamental cause of a problem. Also known as the basic problem, as opposed to a symptom.
A graph in which one node is designated as the root node (starting point) for a search.
A set of column values describing one logical instance in a relational database table. Technically called a tuple in relational calculus. Equivalent to a record in a flat file.
A public key encryption program, named for the authors (Ron Rivest, Adi Shamir, and Leonard Adleman). See also encryption.
The movement of a line or object with one point held fixed while the rest of the object stretches or compresses around that point as other points are moved.
A statement that applies logic or an algorithm to information values to determine a resulting output or action, or to constrain the data relation or its valid values.
Criteria used to determine whether a person or software agent has permission to access data or perform a process.
Generally, a formally stated constraint governing the characteristics or behavior of an object or entity, or the relationship between objects or entities, used to control the complexity of the activities of an enterprise.
In data quality, constraints that can be used to validate the contents of a database. The defined characteristics of a database constitute business rules, such characteristics as dependency/optionality, multiplicity/exclusivity, and value set constraints.
In data analysis, a rule that focuses on a specific set of attributes that uniquely identify an entity and identify merge opportunities without taking automatic action.
In data analysis, a rule that identifies and cross-references records that appear to relate to a master record, without updating the content of the cross-referenced record.
In data analysis, a rule that matches records and merges the data from these records into a single, unified, reconciled, and comprehensive record.
Technologies, applications, and practices for the collection, integration, analysis, and presentation of information to help salespeople keep up to date with clients, prospect data, and drive business. In addition to providing metrics for win-loss and sales confidence, SI can present contextually relevant customer and product information.
Generally, a limited part or subsection of something intended to represent the qualities of the whole.
In data analysis, a selected subset of data from a population, used to better understand the entire population. Samples should be representative of the entire population.
To select a subset of data to test in order to deduce patterns, which then can be compared to the whole for accuracy. Sampling typically has shorter processing windows and therefore can be tested more often until a pattern is defined.
An alternative environment that allows read-only connections to production data and can be managed by users to test hypotheses about data or merge production data with user-developed or supplemental data. Typically used to develop new uses of data.
A US law, enacted in 2002, establishing stringent financial reporting and auditing requirements for publicly traded companies doing business in the United States. It was designed to make executives more responsible and accountable for oversight of their companies. The act covers issues such as auditor independence, corporate governance, internal control assessment, and enhanced financial disclosure. The act also covers security of, and access to, computer systems. Other countries have also adopted the principles of Sarbanes-Oxley.
Choosing the first sensible solution rather than examining all alternate solutions before deciding. Combines the ideas of “satisfy” and “suffice.”
An event-based application programming interface for processing XML documents.
The ability to scale to support larger or smaller volumes of data and more or fewer users. The ability to increase or decrease size or capability in cost-effective increments with minimal impact on the unit cost of business and the procurement of additional services.
Generally, an expression of size, volume, or scope; magnitude, expressed as a ratio of the representation to the actual size.
In architecture, to change in size or capability, in accordance with requirements, with minimal effort or resource impact.
In a numeric figure, the number of places to the left of the decimal place. See also precision.
In data management, one complete set of use-case steps from start to completion.
The design of a dynamic process or financial model to support “what if” analysis, predicting outcomes when variables are changed.
Generally, a diagrammatic representation of the structure, framework, or population of instances of something.
In data management, a data structure.
In some database software, a synonym for an instance of a database management system.
In XML, the set of allowable XML tags, usually expressed in DTD or XSD.
A data model diagram that expresses the structure of a data model in graphical terms, depicting types but not including any actual or sample data values (i.e., instances).
The stored physical database definition derived from a set of DDL (Data Definition Language) statements. The database schema contains all the information that defines the logical database and its physical storage.
A variation of a star schema in which the dimension tables are normalized to remove all transitive dependencies.
A set of relational tables representing multidimensional data, comprised of a single, central fact table surrounded by a single level of denormalized dimension tables. Star schemas implement dimensional data structures with denormalized dimensions. Snowflake schemas are an alternative to star schemas, containing at least one dimension normalized at least one level. The star schema and processes for managing them were invented by Ralph Kimball.
A set of XML tag definitions used to define and document XML applications.
Generally, the boundary within which something has control, power, or obligation.
In project management, the definition of the business or technology impacted by a project’s intended work.
A performance-management tool that uses a structured report to help manage an organization’s performance by reporting a standard set of performance measurements against objectives, internal targets and industry benchmarks. See also Balanced Scorecard (BSC).
The activities and costs required to dispose of something and recreate it from scratch, saving nothing from the prior effort.
An iterative, incremental methodology for project management often seen in Agile software development, a type of software engineering.
An information retrieval system designed to search datasets and return results based on input criteria.
The ability to provide differing access to individuals according to the classification of data and the user’s business function, regardless of variations.
A security protocol designed to ensure the secure transmission of payment information over the internet, primarily for credit card transactions.
The prevention of unauthorized access to a database and its data, and to applications that have authorized access to databases.
A SQL statement (command) that specifies data retrieval operations for rows of data in a relational database.
Generally, the features and characteristics used to narrow choices within a larger field. For example, when evaluating product alternatives.
In data queries, the data values used to select records to form a subset of instances/records/rows of a file or table, expressed as a Boolean expression.
A type of neural networks that uses unsupervised learning to produce two-dimensional representations of an input space.
A process that determines the meaning of a sequence of words in a sentence or documents. In artificial intelligence, it helps machines “understand” and respond to user queries in a more humanlike manner.
Data that has meaning beyond its literal value (e.g., concepts such as customer, product, or credit line); a way of thinking about and organizing data that puts meanings and context first to make it easier for both people and machines to understand and use.
Data integration based on semantics, as opposed to structure. See also data integration.
The degree to which data stored in multiple places is semantically equal in value. For example, one database might use the code value F to designate female gender, another might use the code value 1 to designate female gender; these code values are semantically equal because they stand for the same thing.
The concept in translation that seeks to maintain the same meaning in the target language as found in the source language.
A representation of data using business terms to enable ease of understanding and use.
A single source of truth that standardizes business definitions, metrics, and relationships across the organizations.
A business-friendly abstraction that sits between complex data storage systems (such as data lakes or data warehouses) and end-user tools such as business intelligence dashboards. Its main purpose is to translate raw or technical data structures into clear business terms, thereby enabling consistent analytics without requiring deep technical expertise.
An association of meaning to entities and attributes. See also metadata, business.
A search that interprets the meaning of words and phrases. Semantic search algorithms learn from user behavior to deliver more relevant results.
Internet in which all content is tagged with semantic tags defined in published ontologies. Interlinking these ontologies allows software agents to reason about information not directly connected by document creators.
The study of the meaning behind the syntax (signs and symbols) of a language or graphical expression of something. The semantics can only be understood through the syntax, which is like the encoded representation of the semantics. See also syntax.
The branch of linguistics concerned with signs, symbols, syntax, and semantics, and their use in communication.
Generally, the order of things, or an ordering of things, often numbered.
In data management, a database object that generates numbers in order.
A method of retrieving or touching data in some linear order.
A software service that provides standard functions for clients in response to standard messages from clients.
The physical computer hardware from which services are provided.
A software component invoked via a message. The message may come from outside the service’s environment, and the results returned by the service may be delivered outside the service’s environment (to the requesting component on a different platform).
Originally developed by IBM, a standardized model for organizations to guide their transformation to a service-based business model. The Open Group later adopted SIMM as the foundation for the Open Group Service Integration Maturity Model (OSIMM), the industry’s collaborative maturity model for service-oriented architecture (SOA) adoption.
The part of a contract between two parties that outlines the delivery of services within defined timeframes.
A specific, measurable target for the performance and reliability of a service over a defined period. It represents a promise made by a company regarding metrics such as uptime or incident response. SLOs are key components of service management and are used to set clear expectations and monitor performance in delivering services. They are often part of a broader service-level agreement (SLA) between a service provider and a customer.
The ability to determine the existence of problems, diagnose their causes, and repair and/or solve the problems.
An application architecture organized around the use of services, especially web services.
Performing enterprise application integration (EAI) using service-based technology.
The branch of mathematics that studies collections of objects and the manipulation of those sets.
An early document markup language, since superseded by HTML and XML.
The “plan-do-check-act” cycle of continuous improvement developed by Walter Shewhart and popularized by W. Edwards Deming. See also Deming cycle.
The parsing of an XML document into constituent parts to be stored atomically in a relational database.
A Greek letter (∑) that stands for the sum of a group of numbers.
In statistics, a shorthand term for standard deviation. The Greek letter omicron (ό) stands for the standard deviation of an entire population, and the lowercase English letter (s) stands for the standard deviation of a sample set. See also Six Sigma.
The ratio of meaningful data to nonsense within a data stream.
A process in which the degree of similarity between any two records is scored, most often based on weighted approximate matching between a set of attribute values in the two records. If the score is above a specified threshold, the two records are a match and most likely represent the same entity.
An early project to extend HTML with semantic tags, superseded by Resource Description Framework (RDF).
A wrapper specification from the World Wide Web Consortium (W3C) for requests for web services that facilitates interoperability between a broad mixture of programs and platforms.
Describes a system that allows communication between two endpoints in only one direction. See also half duplex; full duplex.
A model that shows the expected operation of a system based solely on the model.
A process of automatically searching for other objects that may need updating based on the update of one object.
A concept in information management that provides one view of agreed-upon data to provide consistency and accuracy.
A model showing evaluation based on one variable.
In data flow diagrams, where data leaves the data flow without any definition of the target. See also source.
The perception of an environment’s state and conditions at a point in time.
A rigorous and disciplined statistical analysis methodology developed by engineer Bill Smith in 1986 to measure and improve a company’s operational performance, practices, and systems.
(lowercase) In data quality, a level of quality in which six standard deviations of a population fall within the upper and lower control limits of quality, allowing no more than 3.4 defects per million parts or transactions. In many organizations, it has come to mean simply a measure of quality near perfection.
A subset of a multidimensional array corresponding to a single value for one or more members of the dimensions not in the subset. See also slice and dice.
A data analysis function provided by multidimensional tools. Typically refers to allowing a user to filter and sort data in multiple ways.
An interface for attaching disk drives to a CPU, usually on a small computer, via an input/output bus.
Service for standard text messaging on cellular devices.
The state of an object, a system, or a collection of attributes regarding a state at a particular point in time.
A form of psychological manipulation used by attackers to exploit human behavior, rather than technical vulnerabilities, to compromise security systems. Unlike traditional hacking, which targets software or hardware, social engineering targets people.
Websites and applications that allow people to create and share their own content or to participate in social networking.
The use of websites, applications, or platforms to interact with other users, share information, and build personal or professional relationships.
Computer programs, including operating systems, utilities, tools, database management systems, and application programs. Intellectual property that imposes semantic meaning on input from humans and devices.
A distribution method for software through a network interface.
The control of changes made to software and documentation of an information system in development and operational maintenance. Source code/component library management and source code version control are each part of software configuration management.
A set of tools that enables development of system modifications or customizations that will be more likely to properly interface or interact with existing system processes.
An information technology research organization at Carnegie Mellon University in Pittsburgh, PA, funded by the US Department of Defense.
A type of storage device that uses nonvolatile memory to store data, meaning it retains information even when powered off. Sometimes called semiconductor storage device, solid-state device, or solid-state disk.
An algorithm developed to index sounds in order to sort or search text with like sounds.
In data management, a specific dataset, metadata set, database, or metadata repository from where data or metadata are available.
In data flow diagrams, where data enters the data flow. See also sink.
A DCMI element in element set Content: the origination of a resource. See also Dublin Core Metadata Initiative (DCMI).
Human-readable procedural or declarative programming statements that can be compiled into equivalent machine-readable code.
The management of change to software instruction sets over time.
Unsolicited electronic mail advertising. Laws prohibit misrepresenting or falsifying the origin or the routing information using an internet address of a third party without permission. From a Monty Python skit in which chanting the name of a canned meat product overwhelms all other speech.
Enables users (human or other) to query a knowledge base via the SPARQL Protocol and RDF Query Language. Results are typically returned in one or more machine-processable formats. Therefore, a SPARQL endpoint is mostly conceived as a machine-friendly interface toward a knowledge base. Both the formulation of the queries and the human-readable presentation of the results should typically be implemented by the calling software and not be done manually by human users.
A Resource Description Framework query language standardized by the World Wide Web Consortium (W3C).
A community with a shared purpose of promoting some specific subject of interest.
The process of dividing an entity or object class into subtypes based on differing attributes, relationships, and behaviors. The resulting subtypes inherit the characteristics of their more generalized supertype. Contrast with generalization.
The formal documentation of requirements, data definitions, and design descriptions to direct further development.
To support or aid, but not lead, another in an effort.
The extent of variation in a set of items. See also standard deviation.
A concept describing the use of spreadsheets to approximate business intelligence applications. Due to the limitations of spreadsheet applications, multiple redundant applications are developed and spread across an organization, making it difficult to impose standards and formal support.
A two-dimensional format for representing and storing information having columns and rows. A spreadsheet can be used to store a relational table or flat file, assuming the columns have headings and the rows represent entity instances. NOTE: Every flat file or table can be represented in a spreadsheet, but not every spreadsheet is a relational table, even though it consists of columns and rows.
A type of software that gathers information about a person or organization without their knowledge through such means as monitors, cameras adware, or cookies. It may be malicious or it may be used by companies to track web browsing by internet users to serve up ads to potential buyers. Also used to describe the USA National Security Agency’s gathering of personal information through internet sites.
A standard language for accessing relational, Open Database Connectivity, Distributed Relational Database Architecture (DRDA), or nonrelational compliant database systems. The dominant database language, used to define, control, manipulate, and access relational data. Originally an abbreviation for Sequel, but later acquired the reverse acronym of “Structured Query Language” when IBM could not trademark the name, explaining why is it still pronounced “sequel” today.
An end-user tool that accepts SQL to be processed against one or more relational databases.
A security technology commonly used to secure server-to-browser transactions.
An organization, person, process, or system that can be affected by a change to a system or process.
A model or example established by authority, custom, or general consent, used in measurement and comparison of quality, value, quantity, or extent.
A widely used standard taxonomy for classifying businesses as defined by the US Department of Labor, replaced by the North American Industry Classification System (NAICS) taxonomy.
A stored, reusable SQL query that can be issued with or without modification as dynamic SQL to the database. Frequently users provide different parameter values to variables in the standard query to deliver different result sets.
Generally, the way something is at a point in time, as described by its attributes. State is something that is, as opposed to something that happens. Opposite of behavior.
In modeling, a stage in the lifecycle of an entity or object class. Transition to a state is triggered by an event. A state is represented by a status code attribute value.
Part of a state transition diagram, which is data-centric, versus a data flow diagram, which is process-centric.
A representation of the various valid states in the lifecycle of an entity or object class. State transition diagrams are valuable supplemental additions to a data model beyond entity-relationship diagrams. See also data flow diagram.
A stored, precompiled SQL query, optimized for accessibility against a particular database design.
The examination of data to see patterns of probability or effects from causes.
Procedures and methods for measuring process quality, identifying unacceptable performance, variance, and taking corrective action.
The careful, responsible management of something entrusted to one’s care on behalf of others. See also data stewardship.
Involving some chance, randomness, or uncertainty. For example, stochastic analysis.
A detailed and specific product type used by inventory control systems.
A technique using colored circles to identify the content of a data attribute. The colors are defined by a set of predefined thresholds. See also scorecard.
A network system that stores data within the network but separate from exterior networks that allow attachment by application servers.
An alliance of institutions that focuses on promoting storage network industry standards.
A US law that addresses voluntary and compelled disclosure of “stored wire and electronic communications and transaction records” held by third-party internet service providers.
A precompiled code routine stored within a database management system.
The application of business intelligence tools to provide metrics to executives, often in conjunction with some formal method of business performance management, to help determine if a corporation is on target for meeting its goals and objectives. Used to support long-term corporate goals and objectives.
A role accountable for data quality within a major subject area, resolution of business rule and data quality issues, and identification of coordinating and operational data stewards. A member of the data governance council. See also executive data steward.
A set of decisions that sets a direction and defines an approach to solving a problem or achieving a goal.
Data that is generated continuously by various sources such as applications, internet of things sensors, log files, and servers. It is processed, analyzed, and acted upon in real time.
A type of analysis that provides companies with both internal and external factors that could affect the long-term success of the company.
A hierarchical classification for identifying relationships between categories.
A data structure made up of hierarchical relationships between entity types. A hierarchical structure is not a tree structure because the parent entity type and child entity type in any hierarchical relationship need not be from the same population (not homogeneous).
A hierarchy of things from the same population, also called tree-based structure. The things could be:
- instances from a population represented by a single type icon representing the population of instances, and a reflexive relationship on that type, or
- types from the set of types defined in a database represented by a tree structure where each node of the tree is a population of instances of the same type.
In the first case, it is the instances that form a tree structure; in the second, it is the types that form a tree structure. The latter is called a hierarchical data structure.
A set of structured hints to be applied to a family of documents to create a particular type of display.
A topic or central idea.
A DCMI element in element set Content: the area of focus of a resource. See also Dublin Core Metadata Initiative (DCMI).
Generally, a discipline or branch of knowledge.
In data modeling, a group of related entities or tables, logically grouped for presentation and analysis as a view to part of a data model.
An individual with extensive knowledge and expertise in a specific area. SMEs provide guidance and insight on specialized topics.
A data resource that is built from data subjects that represent business objects and events in the real world that are of interest to the organization.
A query called within another query.
A style of system interaction and data distribution in which consumer applications or persons indicate their interest in certain kinds of data from certain sources when certain events occur. When the events occur, the consumers receive the data via a message. This approach is an alternative to continual polling or scheduled batch interfaces.
A specialized subset of occurrences of a more general entity type, having one or more additional attributes or relationships not inherent to other occurrences of the entity. See also supertype; generalization; specialization; primitive.
Tables created along commonly used accessibility dimensions to speed query performance, although the redundancies increase the amount of data in the warehouse. See also aggregate data.
A more generalized entity of which some occurrences belong to a more specialized subtype. See also subtype; generalization; specialization; primitive.
The optimal flow of product from site of production through intermediate locations to the site of final use.
The process of extracting and presenting supply chain information to provide measurement, monitoring, forecasting, and management of the chain.
The process of ensuring optimal flow of inputs and outputs.
A modeling technique that assigns points to classes based on the assignment of previous points, and then determines the gap dividing the classes where the gap is furthest from points in both classes. See also predictive modeling.
Tools and applications designed to gather information about the activities of people. The vast majority of computer surveillance involves monitoring data and traffic on the internet.
An XML-based file format for describing images in terms of two-dimensional shapes, usually for interactive or animated graphic applications.
In computer architecture, the “shared everything” approach to parallel computing. Describes a computer system where all resources are shared, including data storage, memory, and processors. Each task may be processed using any shared resource. Growth is achieved by adding more resources, up to the limits of the hardware. Possible bottlenecks include memory bus contention. Contrast with massively parallel processing.
Describes a style of communication in which the requestor waits for a reply.
The rules governing the encoded representation of a set of semantics, using certain constructs, notations, and grammar. See also semantics.
An extension of UML that adds notation for additional resources such as hardware, software, and facilities.
An interacting and interdependent group of component items forming a unified whole to achieve a common purpose.
A dynamic form of visualization that combines causal loop diagrams and stock and flow diagrams to create a simulation of the workings of a system from one point in time to another.
A system that stores the “official” or authoritative version of a data attribute to ensure data integrity.
An information technology (IT) or business professional responsible for identifying, understanding, and specifying business information requirements and system functional requirements; defining business process models; participating in data modeling and information value chain analysis; and defining test strategies and test plans to verify requirements. Systems analysts also serve as liaisons between IT and business units and as facilitators for organizational and cultural change. See also business analyst; business systems analyst.
The phases and activities common to software development projects. Common phases include initiation, concept development, planning, requirements analysis, design, development, integration and testing, implementation, operations and maintenance, and disposition. Developed by James Martin and Clive Finkelstein in 1981.
The “fifth discipline” of a learning organization, which sees problems in the context of the whole system, applications in the context of the entire value chain, and data as a shared, reusable enterprise resource. See also knowledge management (KM).
In data management, a cluster of data attributes or values associated with a population of entities, each of which is described by the same set of attributes. A cluster of one or more columns to represent information about thed entities. Each attribute must be atomic (single valued). See also flat file; relation.
A term coined by Ralph Kimball to describe a data warehouse table with a multi-part key whose purpose is to capture a many-to-many relationship that cannot be accommodated by the natural grain of a single fact table or dimension table. Similar to an associative table, but specific to dimensional modeling.
A table that serves to link two-dimension tables with a many-to-many relationship that cannot be resolved through a fact table.
A table that captures parent–child relationships within a variable-depth or ragged hierarchy to enable efficient traversal.
In a snowflake schema data mart, a second-level dimension table linked to a primary-level dimension table and not to any fact table.
The process of examining all rows of data in a table sequentially.
A table that is a denormalized hierarchical component of another dimension table.
When the data modeler thinks first of tables when developing a data model for a user domain instead of the data to be used.
The knowledge that a person retains in their mind. It is relatively hard to transfer to others and to disseminate widely. Also known as implicit knowledge.
The application of business intelligence tools to analyze business trends by comparing a metric to the same metric from a previous month or year, or to analyze historical data to discover trends that need attention. Used to support short-term business decisions.
A person who acts as liaison between the strategic data stewards and the detail data stewards to ensure that all business and data concerns are addressed.
The process of implementing a portion of an enterprise data warehouse.
Delimiters in a markup language that also contain information. Matched tags are used in pairs, preceding and following text.
The process of adding metadata or labels to data for easier classification, searching, or retrieval.
Generally, a collection of controlled vocabulary terms organized into a structure of parent–child relationships. Each term is in at least one relationship with another term in the taxonomy. Each parent’s relationship with all of its children is of only one type (whole–part, genus–species, or type–instance). The addition of associative relationships creates a thesaurus.
In content management, a vocabulary (the list of terms in a dialect of an organization or community) organized into a hierarchy, generally to find terms easily.
The hierarchical structure for outlining topics. Dewey decimal classification is an example of a taxonomy.
A taxonomy with no relationship between equal categories. An example is a list of countries.
A taxonomy with a tree structure of at least two levels and with bidirectional relationships. An example is geography from continent to address.
A taxonomy with both hierarchical and facet categories. Any two nodes in a network taxonomy link based on their associations. An example is a thesaurus.
A set of standard protocols used to organize data sent across a network.
A leading provider of in-depth, high-quality education and research in business intelligence and data warehousing. Founded in 1995 as The Data Warehousing Institute, it was renamed in 2015 to better represent its changing scope.
The data that technicians need to build, manage, and maintain databases and make the data available to the business.
A manner of accomplishing a task using technical processes, methods, or knowledge
The application of conceptual knowledge to achieve or improve practical goals.
The application of knowledge to improve performance or productivity, conserve resources, or increase human comfort.
One of the DAMA-DMBOK Functional Framework Environmental Elements. Includes supporting technology (primarily software tools), standards and protocols, product selection criteria, and common learning curves.
A preexisting form or outline that serves as a pattern guideline for creating a document, specification, or software object.
The process of gathering terms in common use which will become the basis of a conceptual model or vocabulary.
Consisting of three components or values.
A validation process that compares in an organized fashion the functionality or content of a thing or process against preestablished requirements for that thing or process.
A set of documented conditions reflecting user requirements.
A dataset that has been specifically created to enable testing of some process using the dataset as a standard input.
A validation process that evaluates functionality between individual components or modules.
A validation process that evaluates the time of system performance during specific activities compared to expected performance parameters.
Retesting existing code using cases of tests that were passed to verify that nothing changed.
A validation process that evaluates hardware and/or software on a complete integrated platform to evaluate compliance with requirements.
A validation process that evaluates functionality of individual code sets or modules, independent of any other code set or module, using a defined set of data.
A validation process that evaluates functionality from a user’s point of view, independent of any technical validation.
The application of statistical, linguistic, and machine learning techniques on text-based data sources to derive meaning or insight.
The process of evaluating unstructured text for patterns to extract actionable data and sentiment via semantic analysis, statistical methods, etc.
A controlled vocabulary with both parent–child and associative relationships defined. See also taxonomy.
A high-level language for manipulating a database one record at a time, with nonessential, repetitive commands encapsulated to make programming more readable by a human. See also fourth-generation language (4GL).
A situation where a large number of resources are involved in doing minimal amounts of work, mostly due to collisions or contention for resources in database access.
A method of controlling the rate at which data is processed or transferred to prevent system overload.
A level of separation of computing responsibility. Originally, computer architecture was monolithic, with all processing occurring on the same machine. Over time, two-tier and three-tier systems separated processing for user interfaces, application logic, and data persistence. Current architecture is n-tiered.
A file format for storing images as electronic files.
A fixed amount of time allocated to a project, activity, task, or meeting.
In project management, a technique for separating parts of a project schedule in order to distribute work, as well as management of that work.
A sequence of data points that can be related to points in time in some pattern of intervals.
The time period when a data value is stored in a database.
The time period when a data value represents a true status in the real world.
The degree to which available data meets the currency requirements of information consumers.
The length of time between data availability and the event or phenomenon described.
A sequence of characters indicating the date and time when a data event occurred.
A system that responds only to a moment in time at which the input is applied.
A term, coined by Malcolm Gladwell, describing the point at which a previously rare phenomenon begins to occur at an epidemic rate.
An identification assigned to an object.
A DCMI element in element set Content: the name of a resource. See also Dublin Core Metadata Initiative (DCMI).
A discrete collection of identifying information.
In operating systems, a container for security information about a user.
A unit of text (word, sub-word, or character) used in natural language processing.
The cost to own, implement, and maintain a product throughout its life.
Techniques, methods, and management principles for continuous improvement, based on the work of W. Edwards Deming, Joseph Juran, Phillip Crosby, and others.
Capable of being related to steps in a process.
In a networked system, the number of packets traversing a network segment.
The process of teaching a machine learning model by exposing it to data and allowing it to adjust its internal parameters.
The algorithm used for training a deep learning system or a predictive analytics model. It figures out what nodes to keep to obtain a good generalization for the problem at hand.
A collection of data whose purpose is to be analyzed to discover patterns that can then be applied to other datasets, or the part of the dataset that is used for training a predictive analytics model before it is tested and deployed.
In business, an event involving the exchange of products, money, and/or data.
In systems, a unit of work including one or more actions performed together or not at all, usually in support of a business transaction.
In databases, a complete atomic unit of work; a set of statements to perform CRUD operations on data, in which the database management system must either complete performance of all the statements or none of the statements. As the process continues, it requests locks on various database objects, according to the concurrent update protocols, deadlock handling scheme, and backup scheme. Any database updates performed are held in limbo until the END transaction statement is encountered. At that point, the system checks the integrity rules to ensure that the database remains in a valid state relative to its definition. If the check shows no errors, the updates are made permanent and all locks on data are released. Otherwise, no changes are applied and the system is reset to the transaction starting point.
An information system designed to store and record day-to-day business information, often structured around events, business processes, or business activities. These systems are optimized for storing large volumes of data and processing a high volume of requests for small amounts of data, but not for analyzing or aggregating data.
In UML, a term for the change of an object’s status from a source state to a target state.
Data that does not exist past the execution of a particular program.
In logic, a relationship where if A and B, and B and C, then A and C.
In data ethics, the characteristic of an organization or program that makes available for public review and comment as much data as possible.
Availability of the amount of data required for reasonable people to make an informed decision about policies and practices that impact them and society as a whole.
In artificial intelligence, the clarity and openness provided by developers about how their models operate, make decisions, or derive conclusions.
Reorganizing data by switching rows and columns, often used in data preparation or analysis.
A graph in which child nodes do not have more than one parent. See also chart; graph; structure, tree.
A predictive analytical technique that can use a combination of continuous and categorical data to produce a decision tree optimized for minimal complexity. CART is related to algorithms such as C4.5 and chi-square automatic interaction detection (CHAID).
A long-term movement in an ordered series (e.g., a time series) that may be regarded (together with the oscillation and random component) as generating the observed values.
A software routine guaranteed to execute when an event occurs. Often a trigger will monitor changes to data values. A trigger includes a monitoring procedure, a set or range of values to check data integrity, and one or more procedures invoked in response, which may update other data or fulfill a data subscription.
A sequence of three entities that reduces a statement about semantic data to machine-readable code in a subject-verb-object term. The “atom” of the Semantic Web. Also called semantic triple.
The direction from any location that points toward the geographic North Pole. Not the same as magnetic north.
The formal mathematical term for a row in a relational table or record instance in a flat file.
A transaction processing protocol that ensures the transaction holds locks on all records involved before committing any updates.
A sampling method that combines samples from a set of common groups and then takes samples from the result.
Generally, a subdivision or category.
In data management, a population of instances defined by a common schema.
A DCMI element in element set Content: the classification of a resource. See also Dublin Core Metadata Initiative (DCMI).
A standard format for an XML registry, an online “yellow pages” directory that gives organizations a uniform way to describe their application services, discover other organizations’ services, and understand the methods required to conduct e-business with a specific company.
The dominant modeling language for object-oriented analysis and design, developed by Ivar Jacobsen, Grady Booch, and James Rumbaugh by consolidating several earlier object-oriented modeling standards. Fundamentally, UML is a process-centric modeling scheme.
Consisting of only one component or value.
In object–role models, describes a predicate consisting of a single object.
To roll back or revert a transaction prior to any commit of that transaction.
A character encoding standard, maintained by the Unicode Consortium, designed to support the use of text in all of the world’s writing systems that can be digitized. Also called the Unicode Standard.
A redundant synonym of identifier. Identifiers are unique by definition. See identifier (ID).
An operation within a unit of work that might need to be undone.
A set of operations performing a logical outcome in which either all changes are successfully performed or none are performed.
A standard coding system used for classifying products and services offered globally.
The US government mail agency, responsible for assigning two-character codes for US states and territories. Similar codes were adopted by Canada for its provinces and territories.
A standard set of characters that may be encoded by different mechanisms.
An agency of the United Nations, responsible for standards in address representation and naming within and for each country.
A system of identifying things by using a sequence of vertical lines of variable thickness. UPCs are scannable representations of 12-digit numeric codes that represent products and their producers.
A type of machine learning where the model is trained on data without labeled responses, aiming to discover hidden patterns or groupings.
Generally, any change to a database, which may also include inserting and deleting data.
In SQL, the change of attribute values for one or more existing rows defined by predicate logic. A SQL statement (command) that specifies replacement of data in a relational database.
A modeling technique that shows the change in probability of an outcome caused by events or actions. See also predictive modeling.
The unambiguous location of a resource in Resource Description Framework. See also Semantic Web.
A string of characters used to identify a resource on the internet. URIs can be either URLs, which provide a means of locating a resource (like a web-page address), or Uniform Resource Names (URNs), which name a resource but do not specify how to locate it.
A resolvable location for a document on the internet. The path information in an HTML-coded source file used to locate another document or image. The format for the URL is scheme://host- domain[:port]/path/filename.
Generally, a description of behavior, given specific input.
In systems development lifecycle, the context for process or workflow scenarios.
In object-oriented analysis, a workflow scenario defined in order to identify objects, their data, and their methods (process steps).
A person or role recognized and authorized to access a particular application. The user’s identity is what confers security authorization. The term has many inappropriate connotations and should be avoided in favor of role-based terms, such as business professional, knowledge worker, data producer, or data consumer.
A colloquial way of describing a user interface that is nonintuitive or difficult to use.
The front end through which users interact with the data.
A method of estimating monetary value of benefit realized by an improvement in worker productivity.
Determining and confirming that something satisfies or conforms to defined rules, business rules, integrity constraints, defined standards, etc. The system cannot perform any validating unless it first has a definition of the way things should be.
The degree to which data conforms to domain values and defined business rules.
Generally, the amount or extent of a measurement of space, time, or quantity.
In data modeling, a data abstraction assigned to a single attribute representing a fact, which may be represented by an encoding of the value.
Commonly, the relative worth, usefulness, desirability, or importance of something, expressed numerically, sometimes using monetary values.
An end-to-end set of activities in support of customer needs, usually beginning with a customer request and ending with customer receipt of benefits. Also called value stream. See also information value chain analysis; process.
The amount of difference or distance between an expected result and an observed result.
The degree to which something is believed to be true.
In language, a word expressing a process or activity that occurs, has occurred, or will occur.
In data modeling, a predicate that implies a relationship between two (or more) objects or entities. It can be used to form an elementary fact sentence.
A specific modification of a basic object that shares the same identifier with other modified forms of the object. Supplementary version numbers or effective dates are often needed to uniquely identify an instance.
The time it takes to redisplay updates to a screen: 1/60 of a second. Used in comparison to zero latency.
Databases that pose unusual performance challenges due to their exceptional size.
The process of combining transistor-based circuits to create integrated chips.
A presentation of a set of data from one or more physical tables as one logical table. A view can include some or all of the rows and columns from each contributing table, and can be defined as the result table from a SELECT statement.
A disk file storage management system invented by International Business Machines Corporation (IBM).
The act of creating a virtual version of something, including computer hardware platforms, storage devices, and computer network resources.
The use of multiple types of visualization formats in one display.
Visual representation of qualitative concepts.
Visual representation of quantitative data in schematic form.
An interactive visual representation of data by transformation into an image that can be then manipulated.
The positioning of information graphically using a secondary related framework to convey insight.
Use of complementary visual representations to enable development and communication of a strategy.
A collection of terms and concepts that have not necessarily been screened for duplication and ambiguity. See also controlled vocabulary.
A way to improve the effectiveness of information storage and retrieval systems, web navigation systems, and other environments that seek to both identify and locate desired content via some sort of description-using language. The primary purpose of vocabulary control is to achieve consistency in the description of content objects and to facilitate retrieval.
The ability of a machine to identify and process human voice input, often used in virtual assistants such as Siri (Apple, Inc.) or Alexa (Amazon.com, Inc.).
Subject to sharp, frequent, or regular changes.
Transient, not persistent.
A tag language to describe a three-dimensional space in a tagged document. Originally called Virtual Reality Markup Language.
See information directory.
See chart, waterfall.
A traditional project management methodology where processes are sequential; often applied in data management projects, data modeling, and extract-transform-load development.
A security measure used to monitor and filter traffic between a web application and the internet, often protecting data from threats.
A system for creating and managing HTML content and other associated web materials.
A program that browses the internet, looking for publicly available resources that can be added to a database for future searching applications, for automation of simple monitoring and maintenance tasks, or for harvesting of specific types of data, such as email addresses. Also called ant, automatic indexer, bot, web spider, web robot, web scutter.
A standard interface for managing interactions with geographic datasets.
The process of discovering useful information and patterns from data on the web.
The process of extracting data from websites using automated tools or scripts.
An OASIS specification for web services to deliver data to internet portals.
An open industry effort chartered to promote web services interoperability across platforms, applications, and programming languages. A diverse community of web services leaders providing guidance, recommended practices, and supporting resources for developing interoperable web services.
A network that connects multiple smaller (local area) networks over a large geographic area.
A symbol, typically an asterisk or question mark, used in data queries or search operations to represent one or more characters. It is often used in databases or search engines to retrieve data patterns.
Knowledge in context; knowledge accumulated and applied in the course of actions. Deep understanding, keen discernment, and a capacity for sound judgment.
The process of passing information verbally between people.
An open-source semantic lexicon for the English language. It groups English words into sets of synonyms called synsets; provides short, general definitions; and records the various semantic relations between these synonym sets. The purpose is twofold: to produce a combination of dictionary and thesaurus that is more intuitively usable, and to support automatic text analysis and artificial intelligence applications. The database and software tools can be downloaded and used freely, and the database can also be browsed online. WordNet was created and is being maintained at Princeton University under the direction of psychology professor George A. Miller. Development began in 1985.
The nonprofit organization responsible for creating and maintaining specifications for interoperable standards on which the World Wide Web is based.
The process of sending updated data or changes back to the source system or source database after it has been processed or analyzed.
Data storage where once the data is written, it cannot be changed and is used as-is.
A specification of the XML needed to invoke a web service listed in a UDDI directory. WSDL provides interface/implementation details of available web services and UDDI registrants. It leverages XML to describe data types, details, interface, location, and protocols. Pronounced “whiz-dull.”
A term describing the situation where a representation or recreation of a thing resembles the original to a close degree. Pronounced “whiz-ee-wig.”
An XML-based markup language developed for financial reporting. It provides a standards-based method to prepare, publish (in a variety of formats), reliably extract, and automatically exchange financial statements according to Generally Accepted Accounting Principles (GAAP) standards. From Extensible Business Reporting Language.
A specification for describing data in a machine-independent format. It ensures that data can be transferred between systems that have different internal representations.
An open XML specification for defining and sharing faceted classifications. From Exchangeable Faceted Metadata Language.
The file format used by Microsoft Excel for storing spreadsheet data. It is often encountered in data management contexts where data is stored and analyzed in spreadsheet format.
A tag-based markup language (a subset of SGML (Standard Generalized Markup Language)) containing a set of rules for encoding documents in machine-readable form, defined by the World Wide Web Consortium (W3C) with a tag set that can be extended. The tags enable XML documents to be self-describing data structures. From Extensible Markup Language.
A set of XML message interfaces using Simple Object Access Protocol (SOAP) to define the data access interaction between a client application and an analytical data provider (online analytical processing (OLAP) and data mining) over the internet. The jointly published XML/A specification allows developers, vendors, and others to query analytical data providers in a standard way. The goal is to provide an open, standard access application program interface for OLAP providers and consumers
The schema language for XML, expressed in XML, equivalent to document type definition in older SGML-family markup languages.
A distribution of data elements where the values or frequency of the data points follow a Y-shaped pattern.
The partitioning of data across different clusters or nodes in distributed systems
Used in reference to application maintenance projects enabling legacy systems to support processing in the new millennium by eliminating ambiguity about century years in dates. This ambiguity was due to hard-coded expressions for the year in software and in databases as two digits (e.g., 1950 was expressed as 50). At the turn of the millennium, accurate calculations would not be possible without proper full four-digit references.
A two-dimensional framework for classifying design artifacts contributing to an enterprise information systems architecture. A grid model based on six basic questions (What, How, Where, Who, When, Why) asked of six stakeholder groups (Planner, Owner, Designer, Builder, Subcontractor, System) to give a holistic view of the enterprise. Conceived by John Zachman in the 1980s.
A data center that runs exclusively on renewable energy to minimize its carbon footprint to zero.
A software vulnerability that is unknown to the vendor or developer and consequently has no fix at the time it is discovered.
A cryptographic method used to prove that something is true without revealing any information about it. It is commonly used to secure data transactions and in transactions that require privacy.
The characteristic of a relationship in which a member of population A may be related to only one member of population B, and a member of population B may not be related to a member of population A. For example, a person (B) and a date of death (A). See also cardinality.
The characteristic of a relationship in which a member of population A may be related to one or more members of population B, and a member of population B may not be related to a member of population A, or may only be related to one member of population A. See also cardinality.
The characteristic of a relationship in which a member of population A may be related to none, one, or multiple members of population B, and a member of population B may not be related to a member of population A, or may only be related to one member of population A. For example, a person (B) and a doctor (A). A person may not have a doctor, and a doctor may have zero, one, or more people as patients. See also cardinality.
A unit of digital information or storage equal to one sextillion bytes. Commonly used to refer to an extremely large volume of data rather than a precise measurement.
The process of organizing storage devices into different zones or logical grouping for better data management, security, or performance.
No terms found
Try a different keyword or browse by letter.
About the DAMA Glossary
Introduction and dedication
Introduction
DAMA International is pleased to present the DAMA Glossary for Data Management and trust that it will serve as a valuable resource for professionals and organizations seeking to improve the use, management, and governance of data.
In this increasingly data-driven world, data management depends upon a common understanding of its concepts, practices, principles, and professional language. To assist in the adoption of these concepts, the DAMA Glossary for Data Management will be not a static publication but an ever-evolving, trusted data reference.
While the core principles of data management remain enduring, the context in which they are applied continues to evolve rapidly as organizations adopt new operating models. Artificial intelligence and machine learning systems increasingly depend on trusted, governed, and high-quality data. At the same time, growing concerns regarding privacy, regulations, ethics, security, and transparency have elevated the importance of responsible data stewardship.
Organizations depend upon data to drive decision-making, enable digital transformation, power artificial intelligence, manage risk, and create sustainable competitive advantage. As a living resource, the DAMA Glossary will benefit from ongoing review, new contributions, and refinement by practitioners and subject matter experts across the global DAMA community. It will continue to build upon decades of professional practices and collective data expertise to offer common, standardized data definitions for the vocabulary used to manage data, including the disciplines of governance, analytics, privacy, master and metadata, data architecture, data quality, and artificial intelligence
DAMA International extends its sincere appreciation to the volunteers, authors, reviewers, working groups, and industry experts whose expertise and dedication have contributed to this publication. Their collective efforts strengthen the profession and help foster a shared understanding of the language – enabling organizations to derive value from data responsibly and effectively.
Dedication
In 1985, when John Zachman founded DAMA, his vision was to have Data Management acknowledged as a recognized profession not unlike architecture or engineering. To do so, he created a universal framework for understanding enterprises, which became the foundation for DAMA and led to his vision being accepted by organizations throughout the globe. To help further his vision, the DAMA Glossary for Data Management provides a shared and consistent vocabulary for effective communication among business leaders, data professionals, technologists, regulators, educators, and researchers.
Share Your Feedback
Help us improve the DAMA glossary
Thank you!
Your submission has been received. Our editorial team reviews all feedback regularly.
See business intelligence, social.