serhii.net

In the middle of the desert you can say anything you want

UNLISTED

06 Dec 2022

Digitalethik reading reflection on paperclips

While I wouldn’t consider myself a rationalist, I’m familiar enough with LessWrong and LW-adjacent communities (and Slate Star Codex / Astral Codex Ten is my favourite blog) . I’ve read some of Eliezer Yudkowsky’s sequences/writings about AI safety and was familiar both with the concept and the game.

AI safety is something that I passionately believe isn’t getting the attention it deserves - somehow all logic goes out of the window once the things start feeling too much “science fiction fantasy”. It’s acceptable to talk about global warming, but an artificial general intelligence getting out of control is hard to take seriously - and by extent, people doing work in this direction have a tiny bit less respect and understanding than people working with on “REAL” dangers.

In Bigfoot lore (I’m not a believer, but reading stuff completely detached from my reality is my favourite way to relax) you’ll discover that a lot of people who had what they thought was a sighting won’t ever talk about this under their real identity - openly talking about that is a good way to lose your job and never be taken seriously ever agan by anyone who googles your name. And my issue with this is that we have literally no way to discuss whether Bigfoot is real.

Back to AGI - the paperclip maximizer is a wonderful game that introduces the concept in a really nice way. It brings home the idea that an AI shouldn’t be EVIL to destroy the universe and do something hard to foresee. It can just do it’s job, and without the relevant safeguards it can just destroy humanity because humanity is the main thing that blocks it from doing it’s job well. Robots aren’t evil like Terminator and therefore science-fiction-y, they don’t even “not have feelings”, such things just aren’t applicable to them - they do what you tell them to do, and once they get powerful enough just pray that your wish is OK. Like a genie from a bottle - be careful what you wish for, because it will be done and done literally and it’s not the job of the genie to think about third-order consequences or to think whether your wish is what what you REALLY want or need or not.

And circling back to Digitalethik - “thinking about consequences” - AI in this context is an example of an actor blindly doing what you tell it to, without asking itself questions like “is this a good idea”.

My papers:

  • How Do Users Discover New Tools in Software Development and Beyond? | SpringerLink:

    Software users rely on software tools such as browser tab controls and spell checkers to work effectively and efficiently, but it is difficult for users to be aware of all the tools that might be useful to them. While there are several potential technical solutions to this difficulty, we know little about social solutions, such as one user telling a peer about a tool. To explore these social solutions, we conducted two studies, an interview study and a diary study. The interview study describes a series of interviews with 18 programmers in industry to explore how tool discovery takes place. To broaden our findings to a wider group of software users, we then conducted a diary study of 76 software users in their workplaces. One finding was that social learning of software tools, while sometimes effective, is infrequent; software users appear to discover tools from peers only once every few months. We describe several implications of our findings, such as that discovery from peers can be enhanced by improving software users’ ability to communicate openly and concisely about tools.

  • Does Transparency in Moderation Really Matter?: User Behavior After Content Removal Explanations on Reddit: Proceedings of the ACM on Human-Computer Interaction: Vol 3, No CSCW

    When posts are removed on a social media platform, users may or may not receive an explanation. What kinds of explanations are provided? Do those explanations matter? Using a sample of 32 million Reddit posts, we characterize the removal explanations that are provided to Redditors, and link them to measures of subsequent user behaviors—including future post submissions and future post removals. Adopting a topic modeling approach, we show that removal explanations often provide information that educate users about the social norms of the community, thereby (theoretically) preparing them to become a productive member. We build regression models that show evidence of removal explanations playing a role in future user activity. Most importantly, we show that offering explanations for content moderation reduces the odds of future post removals. Additionally, explanations provided by human moderators did not have a significant advantage over explanations provided by bots for reducing future post removals. We propose design solutions that can promote the efficient use of explanation mechanisms, reflecting on how automated moderation tools can contribute to this space. Overall, our findings suggest that removal explanations may be under-utilized in moderation practices, and it is potentially worthwhile for community managers to invest time and resources into providing them.

Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus