From mboxrd@z Thu Jan 1 00:00:00 1970 Path: news.gmane.io!.POSTED.blaine.gmane.org!not-for-mail From: Heime Newsgroups: gmane.emacs.help Subject: Re: Regexp capturing unicode characters Date: Wed, 31 Jul 2024 21:50:37 +0000 Message-ID: <7dt77KwxEpd-n-H3ZDksh1CMkQfH10cnniy6lS3hK4lXbwVnyLKQAM602dMs2KNgvINoReUbaEpTzZG3ZlL1n8KTqEGfVk2aMtujCFOjjhs=@protonmail.com> References: Mime-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: quoted-printable Injection-Info: ciao.gmane.io; posting-host="blaine.gmane.org:116.202.254.214"; logging-data="36946"; mail-complaints-to="usenet@ciao.gmane.io" Cc: Heime via Users list for the GNU Emacs text editor To: Heime Original-X-From: help-gnu-emacs-bounces+geh-help-gnu-emacs=m.gmane-mx.org@gnu.org Wed Jul 31 23:52:04 2024 Return-path: Envelope-to: geh-help-gnu-emacs@m.gmane-mx.org Original-Received: from lists.gnu.org ([209.51.188.17]) by ciao.gmane.io with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1sZHEo-0009J6-OE for geh-help-gnu-emacs@m.gmane-mx.org; Wed, 31 Jul 2024 23:52:02 +0200 Original-Received: from localhost ([::1] helo=lists1p.gnu.org) by lists.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1sZHDr-0008U5-0D; Wed, 31 Jul 2024 17:51:03 -0400 Original-Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1sZHDo-0008Rm-2A for help-gnu-emacs@gnu.org; Wed, 31 Jul 2024 17:51:00 -0400 Original-Received: from mail-4319.protonmail.ch ([185.70.43.19]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1sZHDl-0005pV-Us for help-gnu-emacs@gnu.org; Wed, 31 Jul 2024 17:50:59 -0400 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=protonmail.com; s=protonmail3; t=1722462639; x=1722721839; bh=G96n8GsPGa1kZP2cz+mxaFYVmh2eU8yVVzc+KoHXUIA=; h=Date:To:From:Cc:Subject:Message-ID:In-Reply-To:References: Feedback-ID:From:To:Cc:Date:Subject:Reply-To:Feedback-ID: Message-ID:BIMI-Selector; b=XVcB2xQY4MX4ixkJmt9SzD9g41l5YtmRBwdyjhkM29WqW9s/2uooh/yRO47M9GOUL IruhpADUol5UMj5MyeRbRXUg0Upd0zzFMpDV2iFIxLPJL5aCAmkZtqYb7lY/e2kWzS zLUoLthMyEimw2WcXWGg+5jPg2dpAuJa6E7XDh8L5h4J86kfDjI02r18r3vIjC7dxM mUfS2jvhIFoRuxMzJAX2owcYVHDaBZcduOmvffHyvdLSxBeUFtQjf0Oe+bn15x+ZHV YL0LehkRkdHAnSByJdn1ZI9nwUjpKi+YAm4+zyidbIFxB//i5zI8E5G2kMebnLzZEJ stTuHYrYesKHQ== In-Reply-To: Feedback-ID: 57735886:user:proton X-Pm-Message-ID: 0a77f8916249bdade48bb04dde651ecee138784b Received-SPF: pass client-ip=185.70.43.19; envelope-from=heimeborgia@protonmail.com; helo=mail-4319.protonmail.ch X-Spam_score_int: -27 X-Spam_score: -2.8 X-Spam_bar: -- X-Spam_report: (-2.8 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_LOW=-0.7, RCVD_IN_MSPIKE_H4=0.001, RCVD_IN_MSPIKE_WL=0.001, SPF_HELO_PASS=-0.001, SPF_PASS=-0.001, TO_EQ_FM_DIRECT_MX=0.001 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: help-gnu-emacs@gnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: Users list for the GNU Emacs text editor List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: help-gnu-emacs-bounces+geh-help-gnu-emacs=m.gmane-mx.org@gnu.org Original-Sender: help-gnu-emacs-bounces+geh-help-gnu-emacs=m.gmane-mx.org@gnu.org Xref: news.gmane.io gmane.emacs.help:147472 Archived-At: Sent with Proton Mail secure email. On Thursday, August 1st, 2024 at 9:24 AM, Heime wrote: > I am using unicode characters in my elisp code (e.g. foreign language sym= bols in icelandic > and spanish). >=20 > Is the regexp [[:word:]] appropriate to capture them ? >=20 Although I have tried "[:multibyte:]", it did not get matches that I=20 get with "[:word:]". I am using the regexp for constructing imenu expressions. To match=20 ;; DN [g=C5=82=C3=B3wny] Sgn(Major), Lexik(Polish). ("Denotes" ,(concat "^;;\\s-+" "\\([[:word:]]+\\)\\s-+" "\\[\\([[:multibyte:]]+\\)\\]\\s-+") 2)