Data Science: Concepts and PracticeLearn the basics of Data Science through an easy to understand conceptual framework and immediately practice using RapidMiner platform. Whether you are brand new to data science or working on your tenth project, this book will show you how to analyze data, uncover hidden patterns and relationships to aid important decisions and predictions. Data Science has become an essential tool to extract value from data for any organization that collects, stores and processes data as part of its operations. This book is ideal for business users, data analysts, business analysts, engineers, and analytics professionals and for anyone who works with data. You'll be able to: - Gain the necessary knowledge of different data science techniques to extract value from data. - Master the concepts and inner workings of 30 commonly used powerful data science algorithms. - Implement step-by-step data science process using using RapidMiner, an open source GUI based data science platform Data Science techniques covered: Exploratory data analysis, Visualization, Decision trees, Rule induction, k-nearest neighbors, Naïve Bayesian classifiers, Artificial neural networks, Deep learning, Support vector machines, Ensemble models, Random forests, Regression, Recommendation engines, Association analysis, K-Means and Density based clustering, Self organizing maps, Text mining, Time series forecasting, Anomaly detection, Feature selection and more... - Contains fully updated content on data science, including tactics on how to mine business data for information - Presents simple explanations for over twenty powerful data science techniques - Enables the practical use of data science algorithms without the need for programming - Demonstrates processes with practical use cases - Introduces each algorithm or technique and explains the workings of a data science algorithm in plain language - Describes the commonly used setup options for the open source tool RapidMiner |
Contents
| 1 | |
| 19 | |
| 39 | |
4 Classification | 65 |
5 Regression Methods | 165 |
6 Association Analysis | 199 |
7 Clustering | 221 |
8 Model Evaluation | 263 |
12 Time Series Forecasting | 395 |
13 Anomaly Detection | 447 |
14 Feature Selection | 467 |
15 Getting Started with RapidMiner | 491 |
Comparison of Data Science Algorithms | 523 |
About the Authors | 531 |
| 533 | |
Praise | 545 |
Other editions - View all
Common terms and phrases
ARIMA association analysis average base models calculated centroid Chapter chart class label classification coefficients collaborative filtering columns computed correlation credit score data mining data objects data points Data Preparation data science data science process decision tree deep learning default density dimensions distance document ensemble model error evaluation example factors feature selection FIGURE Finance function grid identify implementation input Iris dataset item profile iteration k-NN layer linear regression logistic regression machine learning measure methods movie multiple neural network nodes optimization outlier outlier detection output overfitting parameters petal length predictors Random Forest RapidMiner RapidMiner process ratings matrix recommendation engine regression model relationship rule induction sample scatterplot seasonal series forecasting shown in Fig similar split Sports statistics step supervised learning Support Vector Machine Table target variable test dataset text mining tion training data training dataset training records training set transaction user-item visual weights


