Benjamin-Lee / Benjamin-Lee/deep-rules

If a human can't perform the task with ease, be surprised if a machine can.

Open
#26 11 comments 0 reactions 0 assignees View on GitHub
Tip
Dominant language
HTML
Stars
226
Forks
44
PR merge metrics
No merged PRs in 30d

Description

**Have you checked the [list of proposed rules](https://github.com/Benjamin-Lee/deep-rules/issues?q=is%3Aissue+is%3Aopen+label%3Arule) to see if the rule has already been proposed?**

- [x] Yes

**Feel free to elaborate, rant, and/or ramble.**

Alternative heading: Garbage in -- garbage out.

Some people confuse machine learning -- and deep learning is no exception -- with elaborate statistical methods, and belief that neural networks have magical, super-human powers that can detect patterns nobody else can. As a consequence they belief that they can take their data in any form, throw the trending machine learning method of the day at it (which will automagically find a suitable representation of the data and then learn the task), and then profit.

Life rarely works out that way. Humans are actually incredibly good at detecting patterns. More often than not, if a human can't see "it", a machine won't either (c.f. the rule on the importance on baseline performances). If the machine performs better than chance, be suspicious. It is probably overfitting the training data (c.f. the rule on freezing a test set). For success, it is generally tantamount to find a representation of the data based on which the human could perform the task, and then -- and only then -- try to teach a machine to do it.

**Any citations for the rule?** (peer-reviewed literature preferred but not required)
- [DOI](doi.org/DOI_goes_here)

Contributor guide

Open the contributing guide

Research direction

Review the proposed-rules list linked in the issue and read the existing discussion before deciding whether this rule has a distinct scope. Determine the final rule wording and support it with peer-reviewed literature, replacing the DOI placeholder when the proposal is ready.

Written by the indexing model from the issue text.

Assessment

Tech stack
machine-learning
Domain
documentation
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.