Data Mining with SPSS Modeler: Theory, Exercises and SolutionsNow in its second edition, this textbook introduces readers to the IBM SPSS Modeler and guides them through data mining processes and relevant statistical methods. Focusing on step-by-step tutorials and well-documented examples that help demystify complex mathematical algorithms and computer programs, it also features a variety of exercises and solutions, as well as an accompanying website with data sets and SPSS Modeler streams. While intended for students, the simplicity of the Modeler makes the book useful for anyone wishing to learn about basic and more advanced data mining, and put this knowledge into practice. This revised and updated second edition includes a new chapter on imbalanced data and resampling techniques as well as an extensive case study on the cross-industry standard process for data mining. |
Contents
| 1 | |
| 23 | |
Univariate Statistics | 190 |
Multivariate Statistics | 307 |
Regression Models | 367 |
Factor Analysis | 547 |
Cluster Analysis | 623 |
Classification Models | 753 |
Using R with the Modeler | 1089 |
Imbalanced Data and Resampling Techniques | 1147 |
Case Study Fault Detection in Semiconductor Manufacturing Process | 1192 |
Appendix | 1249 |
Other editions - View all
Data Mining with SPSS Modeler: Theory, Exercises and Solutions Tilo Wendler,Sören Gröttrup No preview available - 2022 |
Data Mining with SPSS Modeler: Theory, Exercises and Solutions Tilo Wendler,Sören Gröttrup No preview available - 2021 |
Common terms and phrases
Analysis node Apply Reset Fig assigned Auto Cluster node calculated Cancel Apply Reset classifier cluster analysis coefficients components cross-validation customers Data Audit node data mining decision boundary define Derive node determine dialog window distance double-click ensemble Euclidean distance example Exercise factor factor analysis Fields Model Expert File Edit File node Filter node final stream function graph input variables K-Means K-Means algorithm leukemia linear regression logistic regression measures method missing values model nugget Neural Network normally distributed number of clusters OK Run Cancel option Ordinal outliers output overfitting parameters Partition node PCA/Factor Plot posttest prediction predictor importance pretest Preview Reclassify records regression model Reset X Fig ROC curve sample scale type scatterplot scores Sect Settings Annotations shown in Fig shows SPSS Modeler standard deviation statistics subset SuperNode Table node target variable template stream test set training data transformed TwoStep Type node Values Clear


