Skip to content

3.12: tokenize adds a newline when it is not there #105259

Description

@asottile

Bug report

tokenize adds a concrete newline character when it is not present. this breaks any sort of roundtrip source generation and pycodestyle's end-of-file checker

here's an example of a file containing a single byte (generated via echo -n 1 > t.py)

# hd -c t.py 
00000000  31                                                |1|
0000000   1                                                            
0000001
# python3.12 -m tokenize t.py
0,0-0,0:            ENCODING       'utf-8'        
1,0-1,1:            NUMBER         '1'            
1,1-1,2:            NEWLINE        '\n'           
2,0-2,0:            ENDMARKER      ''             
# python3.11 -m tokenize t.py
0,0-0,0:            ENCODING       'utf-8'        
1,0-1,1:            NUMBER         '1'            
1,1-1,2:            NEWLINE        ''             
2,0-2,0:            ENDMARKER      '' 

Your environment

  • CPython versions tested on: 3.12 dbd7d7c
  • Operating system and architecture: ubuntu 22.04 LTS x86_64

cc @pablogsal

Linked PRs

Activity

  1. pablogsal commented on Jun 3, 2023

    @pablogsal
    Member

    @lysnikolaou can you take a look at this?

    I think this is due to this line:

    /* Last line does not end in \n, fake one */

    in combination with this line:

    str = PyUnicode_FromString("\n");

    But the real question is what is going to break if we cover the first line with if (!tok->extra_tokens) and how we can fix it.~

  2. pablogsal commented on Jun 3, 2023

    @pablogsal
    Member

    @lysnikolaou The problem is that we are always adding a new line at the end of the line buffer so we cannot distinguish if the last one is there for real or not (also, we always hardcode \n for NEWLINE tokens which we should not do for the last one).

  3. lysnikolaou commented on Jun 5, 2023

    @lysnikolaou
    Member

    The problem is that we are always adding a new line at the end of the line buffer so we cannot distinguish if the last one is there for real or not

    Do we? I thought you had removed this in #104850, but I guess I misunderstood that PR. I'll have a look.

  4. pablogsal commented on Jun 5, 2023

    @pablogsal
    Member

    Do we? I thought you had removed this in #104850, but I guess I misunderstood that PR. I'll have a look.

    That was "reverted" because apparently that was required to keep backward compatibility IIRC. But this is also a different kind of "new line": the tokenize module needs new lines at the end of every logical line to correctly emit the last tokens. Just comment out this part and see what fails (locations of last tokens are wrong as well as the some implicit dedents are missing):

    /* Last line does not end in \n, fake one */

    Notice that just removing that is not enough to fix this: we would also need to emit the last artificial newline if the input doesn't contain it.

  5. pablogsal commented on Jun 6, 2023

    @pablogsal
    Member

    @Yhg1s can you wait until we fix this for the release?

  6. added a commit that references this issue on Jun 6, 2023
  7. Yhg1s commented on Jun 6, 2023

    @Yhg1s
    Member

    Yeah, I can wait.

  8. added 2 commits that reference this issue on Jun 6, 2023
  9. added a commit that references this issue on Jun 6, 2023
  10. added a commit that references this issue on Jun 6, 2023
  11. pablogsal commented on Jun 6, 2023

    @pablogsal
    Member

    @asottile can you confirm this fixes it?

  12. asottile commented on Jun 6, 2023

    @asottile
    ContributorAuthor

    yep looks fixed -- this is the only outstanding breaking change I'm encountering: #105390

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

3.12only security fixes3.13only security fixesrelease-blockertype-bugAn unexpected behavior, bug, or error

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions