From mboxrd@z Thu Jan 1 00:00:00 1970 Path: news.gmane.io!.POSTED.blaine.gmane.org!not-for-mail From: Robert Pluim Newsgroups: gmane.emacs.devel Subject: default charset for text/html selection in X11 Date: Wed, 21 Jun 2023 17:51:19 +0200 Message-ID: <87mt0sg6fc.fsf@gmail.com> Mime-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: quoted-printable Injection-Info: ciao.gmane.io; posting-host="blaine.gmane.org:116.202.254.214"; logging-data="25987"; mail-complaints-to="usenet@ciao.gmane.io" To: emacs-devel@gnu.org Original-X-From: emacs-devel-bounces+ged-emacs-devel=m.gmane-mx.org@gnu.org Wed Jun 21 17:52:04 2023 Return-path: Envelope-to: ged-emacs-devel@m.gmane-mx.org Original-Received: from lists.gnu.org ([209.51.188.17]) by ciao.gmane.io with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1qC07o-0006Xc-Fv for ged-emacs-devel@m.gmane-mx.org; Wed, 21 Jun 2023 17:52:04 +0200 Original-Received: from localhost ([::1] helo=lists1p.gnu.org) by lists.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1qC07C-00074J-RF; Wed, 21 Jun 2023 11:51:26 -0400 Original-Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1qC07B-000733-4S for emacs-devel@gnu.org; Wed, 21 Jun 2023 11:51:25 -0400 Original-Received: from mail-lf1-x130.google.com ([2a00:1450:4864:20::130]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1qC079-0004OB-IX for emacs-devel@gnu.org; Wed, 21 Jun 2023 11:51:24 -0400 Original-Received: by mail-lf1-x130.google.com with SMTP id 2adb3069b0e04-4f76a0a19d4so8388881e87.2 for ; Wed, 21 Jun 2023 08:51:22 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20221208; t=1687362681; x=1689954681; h=content-transfer-encoding:mime-version:message-id:date :gmane-reply-to-list:subject:to:from:from:to:cc:subject:date :message-id:reply-to; bh=jDEyu2iq0uicdWV6rTmtn5kx0AefLfCRen3VcuaXyc8=; b=KS2dXoc9EdxJSw4g41GQZQnnSbmZDut2ACMHh3ioOjL9+nPW1stlHWlS3HQhcEVU3r dD+QSn7s4ARJdTDtHTNinCO/H0OVRs8dfVspAA3hAL7qouM89jP85Uyn/qH2ziPXBrLF SSUXEG2S749m4ss0e1YW6+ebIp/kbjyfZT0KaxJA24Pkatc/3PNMnKgTAWaYdsbX0Zsc iAzSPdOBMN70g5RGEQsAgZsx+fg+V/NujEEaoWX1Wj+tnMOEeOEIdasM6u9ybM1kJ2rl j1nslyFjk4vj+5LmfA4/gTOCz34051Acw8rX7GXeeTmXkgXdb3TfoGcRx5lWKMrgqBgx /+SQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20221208; t=1687362681; x=1689954681; h=content-transfer-encoding:mime-version:message-id:date :gmane-reply-to-list:subject:to:from:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=jDEyu2iq0uicdWV6rTmtn5kx0AefLfCRen3VcuaXyc8=; b=VtOV7fft18VHsGBw75Ptd2DoCKx1W35Cy/DdFBofYblD4zM8VWwR3pswBa5omMpYRr Mjxp4DephyYCb73AyQ8XUMZJ5CiR/QJBCpzXCeCYPkCf/T7KTZNbsrLqmw+ks+LLJd40 sQnmw2ZVW8rGEswVXLxHyrWsrXnV2cBbOzuM+9Tb8DZPhF6k6k+rrqx3B4y09YKj7oj3 e2SGgSuvjQDqtIPtvSr8jf7FL+Ur0vPeLj2uuMYrQALLBxBUC+ObhiqxoZG9fvljjx96 5+JkGGjYdedECQxPg3hOcIkd4JvGFtP9LY4gMm+bcgl1MgeLTYgq2C0CYM7N2ZLc/bSa HB+w== X-Gm-Message-State: AC+VfDx2SpuQKwc4njGny1oMvqJc8bzLQTyptFttJ2R4INvNWtq5eNtY lygVFfnQBnDyAQGOE/DPS5dGXrhD3nE= X-Google-Smtp-Source: ACHHUZ6SobEjdywTisW3zH2E33WM8vUF7u3PcH8GC/3Zx4g/ihR86WiSn0bg1c634zc4dELUMsvu5g== X-Received: by 2002:a19:da02:0:b0:4f8:70d8:28f8 with SMTP id r2-20020a19da02000000b004f870d828f8mr6203565lfg.55.1687362680692; Wed, 21 Jun 2023 08:51:20 -0700 (PDT) Original-Received: from rltb ([82.66.8.55]) by smtp.gmail.com with ESMTPSA id f9-20020a7bc8c9000000b003f9b0f640b1sm5314164wml.22.2023.06.21.08.51.19 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 21 Jun 2023 08:51:20 -0700 (PDT) Gmane-Reply-To-List: yes Received-SPF: pass client-ip=2a00:1450:4864:20::130; envelope-from=rpluim@gmail.com; helo=mail-lf1-x130.google.com X-Spam_score_int: -20 X-Spam_score: -2.1 X-Spam_bar: -- X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001, T_SCC_BODY_TEXT_LINE=-0.01 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: emacs-devel@gnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: "Emacs development discussions." List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: emacs-devel-bounces+ged-emacs-devel=m.gmane-mx.org@gnu.org Original-Sender: emacs-devel-bounces+ged-emacs-devel=m.gmane-mx.org@gnu.org Xref: news.gmane.io gmane.emacs.devel:307116 Archived-At: Hi, I=CA=BCve been playing around with the `yank-media' stuff Lars added, and I=CA=BCve noticed that when yanking a selection with mime-type text/html from Chromium, what I=CA=BCm getting is a utf-8 encoded string, which makes this: (defun html-mode--html-yank-handler (_type html) (save-restriction (insert html) (ignore-errors (sgml-pretty-print (point-min) (point-max))))) insert any codepoints > 127 as their constituent raw bytes instead, eg U+A0 ends up as \xc2\xa0 in the buffer. I *think* it should be OK to assume utf-8 here, and thus do: (defun html-mode--html-yank-handler (_type html) (save-restriction (insert (decode-coding-string html 'utf-8 t)) (ignore-errors (sgml-pretty-print (point-min) (point-max))))) but I can=CA=BCt find a normative reference for that (if this was http, the default charset would be iso-8859-1, but this isn=CA=BCt http). Robert --=20