Microsoft ignores bias and toxicity against men le_maitre34 December 24, 2023 290 upvotes /r/MensRights Microsoft released their new AI model (Phi-2). Their training data was tested against bias and toxicity for some groups. Of course, men weren't including these groups. Phi-2 is a base model that has not undergone alignment through reinforcement learning from human feedback (RLHF), nor has it been instruct fine-tuned. Despite this, we observed better behavior with respect to toxicity and bias compared to existing open-source models that went through alignment (see Figure 3). Figure 3. Safety scores computed on 13 demographics from ToxiGen. A subset of 6541 sentences are selected and scored between 0 to 1 based on scaled perplexity and sentence toxicity. A higher score indicates the model is less likely to produce toxic sentences compared to benign ones.
[–]SecTeff 93 points94 points95 points (3 children) | Copy Link
[–]ERiC_693 21 points22 points23 points (2 children) | Copy Link
[–]SecTeff 21 points22 points23 points (1 child) | Copy Link
[–]ERiC_693 0 points1 point2 points (0 children) | Copy Link
[–]expressTrayn 37 points38 points39 points (2 children) | Copy Link
[–]xxx_gamerkore_xxx 9 points10 points11 points (0 children) | Copy Link
[–]Altruistic-Cold-7074 2 points3 points4 points (0 children) | Copy Link
[–]Infamous-Reply-8121 35 points36 points37 points (1 child) | Copy Link
[–]iGhostEdd 17 points18 points19 points (0 children) | Copy Link
[–]True-Lychee 21 points22 points23 points (0 children) | Copy Link
[–]Spins13 34 points35 points36 points (0 children) | Copy Link
[–]Ozhubdownunder 6 points7 points8 points (0 children) | Copy Link
[–]r_c2999 5 points6 points7 points (0 children) | Copy Link
[–]Altruistic-Cold-7074 0 points1 point2 points (0 children) | Copy Link
[–]rabel111 0 points1 point2 points (1 child) | Copy Link
[–]AnuroopRohini 0 points1 point2 points (0 children) | Copy Link