I'm glad to have kept reading to the author's conclusion: > As a hybrid approach...

ma2rten · on Sept 22, 2018

This approach only works if you use OP's assumption that a text's sentiment is the average of it's word's sentiment. That assumption is obviously flawed (e.g. "The movie was not boring at all" would have negative sentiment).

Making this assumption is fine in some cases (for example if you don't have training data for your domain), but if you build a classifier based on this assumption why don't you just use an off-the-shelf sentiment lexicon? Do you really need to assign a sentiment to every noun known to mankind? I doubt that this improves the classification results regardless of the bias problem.

jakelazaroff · on Sept 22, 2018

Sure, it's flawed, but that's the point of the post: that assumptions about your dataset can lead to unexpected forms of bias.

> Do you really need to assign a sentiment to very noun known to mankind?

No, but it seems like a simple (and seemingly innocuous) mistake that many programmers can and will make.

ma2rten · on Sept 22, 2018

I was just trying to explain in this comment why I think the human moderation solution is solving the wrong problem.

swingline-747 · on Sept 22, 2018

Heck, it's so important that it needs people with detail-orientation and solid judgement, because crowdsourcing (ie populism) may not be the best source of Godwin's law ethical mooring.

User23 · on Sept 22, 2018

The old Wise and Benevolent Philosopher King model of governance applied to machine learning?

rhizome · on Sept 22, 2018

Another point in favor of having moderators.