Kjell Johnson and I are working on a new book: Applied Machine Learning for Tabular Data.

aml4td.org/

Content on pre-processing, feature engineering, supervised models, and post-processing.

We would love to get feedback.

Written with #quartopub and #rstats

No code on the main site but there is a computing supplement using #tidymodels. We hope to have a scikit-learn supplement too.

Follow

@topepo when you talk of tabular data, do you consider that this is only a form that could have very different meanings? E.g. both relational (linked multi-tabular) and CSV (untyped) are tabular. Then, beyond the typical number series stats, there is also ordinal, categorical and nominal data. Would you consider that applying general statistics on such tabular data is implicitly performing transformations that might require quite sophisticated analysis as justification? These are questions relevant to data that you might consider to be tabular, but they are not about the circumstance that it was decided to represent the data in a table.

@mapto a lot of that is spelled out in the introductory chapter. The nomenclature is tough here because there’s no standardized ontology for data that everybody is familiar with.

The reason that “tabular” is in the title is because when people say “machine learning,“ they automatically assume we are talking about image analysis and similar non-tabular data. That wasn’t true 10 years ago, but is very much the case now.

Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.