Knowledge Discovery in the Social Sciences: A Data Mining ApproachKnowledge Discovery in the Social Sciences helps readers find valid, meaningful, and useful information. It is written for researchers and data analysts as well as students who have no prior experience in statistics or computer science. Suitable for a variety of classes—including upper-division courses for undergraduates, introductory courses for graduate students, and courses in data management and advanced statistical methods—the book guides readers in the application of data mining techniques and illustrates the significance of newly discovered knowledge. Readers will learn to: • appreciate the role of data mining in scientific research • develop an understanding of fundamental concepts of data mining and knowledge discovery • use software to carry out data mining tasks • select and assess appropriate models to ensure findings are valid and meaningful • develop basic skills in data preparation, data mining, model selection, and validation • apply concepts with end-of-chapter exercises and review summaries |
Contents
New Contributions and Challenges | 18 |
Data Issues | 43 |
Data Visualization | 70 |
Assessment of Models | 93 |
Cluster Analysis | 115 |
Associations | 133 |
Generalized Regression | 155 |
Other editions - View all
Knowledge Discovery in the Social Sciences: A Data Mining Approach Xiaoling Shu Limited preview - 2020 |
Knowledge Discovery in the Social Sciences: A Data Mining Approach Xiaoling Shu Limited preview - 2020 |
Knowledge Discovery in the Social Sciences: A Data Mining Approach Prof. Xiaoling Shu Limited preview - 2020 |
Common terms and phrases
actors adjacency matrix Age=Adult algorithms allowing anti-religionists analyze approach association rule mining big data Binning box plot Box-Cox transformation calculate causal Class=2nd Class=3rd classification cluster analysis coefficient complex Computational Social Science conventional statistical cross-validation curve data mining data set database decision tree distance distribution documents equation estimate evaluate example expectancy at birth Facebook figure frequency function GDP per capita Gender graph hidden layer identify imputation income input variables itemsets knowledge discovery logistic regression machine learning Manhattan Distance matrix measures missing data multiple n-gram neural networks nodes normal outcome variable output variable overfitting package parameters patterns percent prediction predictors provides Race & Ethnicity Random Forest ratio regression model relationship Research Methods sample Sex=Male sigmoid function skewness split squared errors structure sum of squares supervised learning Survived=No target variable testing set text mining theory tion topics training set transformation unsupervised web mining


