From mboxrd@z Thu Jan 1 00:00:00 1970 Path: news.gmane.io!.POSTED.blaine.gmane.org!not-for-mail From: Yuan Fu Newsgroups: gmane.emacs.devel Subject: Re: treesitter local parser: huge slowdown and memory usage in a long file Date: Mon, 3 Jun 2024 21:53:47 -0700 Message-ID: References: <2DB11528-C657-4AC1-A143-A13B1EAC897A@gmail.com> <0132CFC2-CFA0-4D58-9632-6E6E03FE57DB@gmail.com> <8E3466C4-0875-4187-ADC3-5C72FF23A24F@gmail.com> <81dab46b-dba3-45d0-b509-1d40f4b116bf@gutov.dev> <6D101DD5-6201-4CA6-A105-28A6DA32C3DF@gmail.com> <46b255d5-d8ec-49ce-b649-02ce8488e873@gutov.dev> Mime-Version: 1.0 (Mac OS X Mail 16.0 \(3774.600.62\)) Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: quoted-printable Injection-Info: ciao.gmane.io; posting-host="blaine.gmane.org:116.202.254.214"; logging-data="30693"; mail-complaints-to="usenet@ciao.gmane.io" Cc: "Ergus via Emacs development discussions." , Stefan Monnier To: Dmitry Gutov Original-X-From: emacs-devel-bounces+ged-emacs-devel=m.gmane-mx.org@gnu.org Tue Jun 04 06:55:03 2024 Return-path: Envelope-to: ged-emacs-devel@m.gmane-mx.org Original-Received: from lists.gnu.org ([209.51.188.17]) by ciao.gmane.io with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1sEMCM-0007hx-Jo for ged-emacs-devel@m.gmane-mx.org; Tue, 04 Jun 2024 06:55:02 +0200 Original-Received: from localhost ([::1] helo=lists1p.gnu.org) by lists.gnu.org with esmtp (Exim 4.90_1) (envelope-from ) id 1sEMBQ-0003n5-D7; Tue, 04 Jun 2024 00:54:04 -0400 Original-Received: from eggs.gnu.org ([2001:470:142:3::10]) by lists.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.90_1) (envelope-from ) id 1sEMBO-0003mg-Oz for emacs-devel@gnu.org; Tue, 04 Jun 2024 00:54:02 -0400 Original-Received: from mail-pf1-x433.google.com ([2607:f8b0:4864:20::433]) by eggs.gnu.org with esmtps (TLS1.2:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.90_1) (envelope-from ) id 1sEMBN-0005WH-4k for emacs-devel@gnu.org; Tue, 04 Jun 2024 00:54:02 -0400 Original-Received: by mail-pf1-x433.google.com with SMTP id d2e1a72fcca58-702492172e3so743716b3a.0 for ; Mon, 03 Jun 2024 21:54:00 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1717476839; x=1718081639; darn=gnu.org; h=to:references:message-id:content-transfer-encoding:cc:date :in-reply-to:from:subject:mime-version:from:to:cc:subject:date :message-id:reply-to; bh=dzs+abzEqZxiqpkv1A9v7xrnRXQAfyThRu384GzfxTI=; b=R1TAwoWbU7rS6fs76DSY6y03Rq5nl8nD7OiP5pn3FL96Y3NMLbPXxNhnksD75rlVGq 9lEYB1SxOtCr2/hyqABWHchOeYzy7FbByUSdbhhkSLEWqfik72hp1acF1Q0h8kqfnFmn GIc+SDynbi0mkS2YVmmQ7243iSe5D5o1I2ASebTahBo3RehDGTrm2ZL7hz2+6dsse/y9 X0SQWgur43bLF5kz81D5ZO9mAg91TcQFI7wTpJZWAJ/nvVlM7BLpTOFvRT7FjI+IX8hl fj7BpLLDyivmkrBdPgE+XZPVc7NmwMRQJDs8P9gLijbYHX2qXyDNT3bVd5RmhJ6pJI9M TkMA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1717476839; x=1718081639; h=to:references:message-id:content-transfer-encoding:cc:date :in-reply-to:from:subject:mime-version:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=dzs+abzEqZxiqpkv1A9v7xrnRXQAfyThRu384GzfxTI=; b=j3JsUvZ7dUKhEWy+POeWB1rJeARLPHt7gaX2N3NeQrTGbuzQUpkPQK7ZVJgc9XGM0B 3ldcIkaqKy9lwBZ526mjgUZCAVGfjQvNNskiafoYzMCMVR5EdMqXigLigMk6iVuI1sYT eMIHs6L2Yle0ucd/kDgMhvQ+stMvXsPIXxQavvUB6tUjgbvB4r3frUPGP5s6N+dDAagH OXEQfnt7DJf0oLAXYNb+gjSz0kWM9BJNCPMf4lu4yamPsZNDhxoVQjaB551L/OZLKQ0K +gL2iZVtqmqe3sQYkwS2HgmohvYIv0smzvHN2o9XWbBb32QvyVDEn3TNNvlVvePGYLAT 03Fg== X-Gm-Message-State: AOJu0YyADS24Ie5YSqwFSkXH0wJUTfC82jGYPnmb+X49MSEyDgXH2bm2 7P0ISDER/UncWx4QKqJPmixkCUWZvLn+l+ChNwUYq0LD+PPaIDUBkPplKQ== X-Google-Smtp-Source: AGHT+IHfedN7rRvCE+u33BjbQodlXPOencaoQT6iG7EZrNigvOy1j0+JqmOnpy+pRwTkCeEklB+ZPQ== X-Received: by 2002:a05:6a20:734d:b0:1a9:8152:5102 with SMTP id adf61e73a8af0-1b26f121370mr13851379637.24.1717476839281; Mon, 03 Jun 2024 21:53:59 -0700 (PDT) Original-Received: from smtpclient.apple ([2601:646:8f81:f810:6d47:6e61:bc5f:51a3]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-1f632339066sm75691635ad.22.2024.06.03.21.53.58 (version=TLS1_2 cipher=ECDHE-ECDSA-AES128-GCM-SHA256 bits=128/128); Mon, 03 Jun 2024 21:53:58 -0700 (PDT) In-Reply-To: <46b255d5-d8ec-49ce-b649-02ce8488e873@gutov.dev> X-Mailer: Apple Mail (2.3774.600.62) Received-SPF: pass client-ip=2607:f8b0:4864:20::433; envelope-from=casouri@gmail.com; helo=mail-pf1-x433.google.com X-Spam_score_int: -20 X-Spam_score: -2.1 X-Spam_bar: -- X-Spam_report: (-2.1 / 5.0 requ) BAYES_00=-1.9, DKIM_SIGNED=0.1, DKIM_VALID=-0.1, DKIM_VALID_AU=-0.1, DKIM_VALID_EF=-0.1, FREEMAIL_FROM=0.001, RCVD_IN_DNSWL_NONE=-0.0001, SPF_HELO_NONE=0.001, SPF_PASS=-0.001, T_SCC_BODY_TEXT_LINE=-0.01 autolearn=ham autolearn_force=no X-Spam_action: no action X-BeenThere: emacs-devel@gnu.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: "Emacs development discussions." List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Errors-To: emacs-devel-bounces+ged-emacs-devel=m.gmane-mx.org@gnu.org Original-Sender: emacs-devel-bounces+ged-emacs-devel=m.gmane-mx.org@gnu.org Xref: news.gmane.io gmane.emacs.devel:319809 Archived-At: > On May 27, 2024, at 3:24=E2=80=AFPM, Dmitry Gutov = wrote: >=20 > On 28/05/2024 01:03, Yuan Fu wrote: >=20 >>> But if one operation just changes text in that range (keeping its = length intact, e.g. capitalizing the whole region), and another does the = same (back to lower case), then the combined range would remain = 200..300. >>>=20 >>> Computing that might be difficult without having access to the kinds = of changes are being done (does tree-sitter report those?). OTOH, most = of the time the most important part is the position of the beginning of = the changes (e.g. for syntax-ppss), and we could treat the rest of the = buffer as invalidated=E2=80=A6 >> Oh you=E2=80=99re absolutely right, the range will be shifted by = later edits in the buffer. It=E2=80=99ll be hella hairy to keep track of = all that=E2=80=94say the previous changed range is (100 . 200), and user = inserted 50 chars in position 150, we need to account for that and = update the range to (100 . 250) before merging the new updated ranges = with this one. >> So it seems the best way is really to move treesit--pre-redisplay = entirely into the primary parser=E2=80=99s notifier, WDYT? >=20 > Yep, that sounds easier. And the performance should be about the same, = even if it'd have a bit extra overhead in those theoretical complex = cases. >=20 Ok, I pushed a commit to master that does just that. I tried with C=E2=80=99= s block comment, and php-ts-mode. Everything seems to work fine. I also added treesit-primary-parser. This is supposed to be another = configuration variable that a major mode should set. I=E2=80=99ve = encountered various cases where knowing the primary parser (parser that = parses the entire buffer rather than just a subset of it) would be very = helpful. Treesit-primary-parser can be auto-guessed if major mode = doesn=E2=80=99t set it, so it shouldn=E2=80=99t break anything. I=E2=80=99= d love to know yours and Stefan=E2=80=99s thoughts on it. Yuan=