파이썬에서 문자열 내에서 여러 문자열 찾기


81

파이썬에서 문자열 내에서 여러 문자열을 찾으려면 어떻게해야합니까? 이걸 고려하세요:

>>> text = "Allowed Hello Hollow"
>>> text.find("ll")
1
>>> 

그래서 첫 번째 발생 ll 예상대로 1입니다. 다음 발생을 어떻게 찾습니까?

동일한 질문이 목록에 유효합니다. 중히 여기다:

>>> x = ['ll', 'ok', 'll']

ll인덱스로 모든를 어떻게 찾 습니까?


3
>>> text.count ( "ll")
blackappy

3
이 카운트 발행 수를 @blackappy, 그들을 지역화하지 않습니다
pcko1

답변:


123

정규식을 사용 re.finditer하여 겹치지 않는 모든 항목을 찾을 수 있습니다 .

>>> import re
>>> text = 'Allowed Hello Hollow'
>>> for m in re.finditer('ll', text):
         print('ll found', m.start(), m.end())

ll found 1 3
ll found 10 12
ll found 16 18

또는 정규 표현식의 오버 헤드를 원하지 않는 경우 다음 인덱스 str.find를 가져 오기 위해 반복적으로 사용할 수도 있습니다 .

>>> text = 'Allowed Hello Hollow'
>>> index = 0
>>> while index < len(text):
        index = text.find('ll', index)
        if index == -1:
            break
        print('ll found at', index)
        index += 2 # +2 because len('ll') == 2

ll found at  1
ll found at  10
ll found at  16

이것은 목록 및 기타 시퀀스에도 적용됩니다.


1
정규식을 사용하지 않고 할 수있는 방법이 없습니까?
user225312

1
문제가있는 것이 아니라 단지 궁금합니다.
user225312

2
목록에는 find. 그러나 그것은 작동합니다 index, 당신은 단지 except ValueError-1에 대한 테스트 대신에 필요합니다
aaronasterling

@Aaron : 기본 아이디어를 언급하고 있었는데, 물론 목록에 대해 약간 수정해야합니다 (예 : index += 1대신).
찌르기

4
이제 모든 index += 2것을 언급 했으므로 이것을 문자열 'llll'에 적용하면 'll'의 4 개 중 2 개가 누락됩니다. index += 1현에 대해서도 고수하는 것이 가장 좋습니다.
aaronasterling 2010 년

31

당신이 찾고있는 것은 string.count

"Allowed Hello Hollow".count('ll')
>>> 3

이것이 도움이되기를 바랍니다.
참고 : 이것은 겹치지 않는 발생만을 캡처합니다.


우와, 고마워요. 이것은 매우 간단한 작업 대답했다
데릭

이것은 문자열에서 주어진 부분 문자열의 Count number에 대한 답변 입니다. 일치하는 항목의 인덱스를 찾는 것이 아니라 실제 질문이 아닙니다.
Tomerikoo

26

목록 예의 경우 이해력을 사용하십시오.

>>> l = ['ll', 'xx', 'll']
>>> print [n for (n, e) in enumerate(l) if e == 'll']
[0, 2]

문자열의 경우 :

>>> text = "Allowed Hello Hollow"
>>> print [n for n in xrange(len(text)) if text.find('ll', n) == n]
[1, 10, 16]

이것은 "ll"의 인접 실행을 나열합니다.

>>> text = 'Alllowed Hello Holllow'
>>> print [n for n in xrange(len(text)) if text.find('ll', n) == n]
[1, 2, 11, 17, 18]

와우 좋아. 감사합니다. 이것은 완벽 해요.
user225312

5
이것은 매우 비효율적입니다.
Clément

1
클레망 @보다 효율적 예를 게시
sirvon

@ Clément print [n for n in xrange (len (text)) if text [n-1 : n] == 'll']
Stephen

내 의미 : print [n for n in xrange (len (text)) if text [n : n + 2] == 'll']
Stephen

14

FWIW, 여기에 poke의 솔루션 보다 깔끔하다고 생각하는 비 RE 대안이 몇 가지 있습니다. .

첫 번째 사용 str.index및 확인 ValueError:

def findall(sub, string):
    """
    >>> text = "Allowed Hello Hollow"
    >>> tuple(findall('ll', text))
    (1, 10, 16)
    """
    index = 0 - len(sub)
    try:
        while True:
            index = string.index(sub, index + len(sub))
            yield index
    except ValueError:
        pass

제 용도 테스트 str.find의 센티넬과 검사를 -1사용하여 iter:

def findall_iter(sub, string):
    """
    >>> text = "Allowed Hello Hollow"
    >>> tuple(findall_iter('ll', text))
    (1, 10, 16)
    """
    def next_index(length):
        index = 0 - length
        while True:
            index = string.find(sub, index + length)
            yield index
    return iter(next_index(len(sub)).next, -1)

이러한 함수를 목록, 튜플 또는 기타 반복 가능한 문자열에 적용하려면 다음과 같이 함수를 인수 중 하나로 취하는 상위 수준 함수를 사용할 수 있습니다 .

def findall_each(findall, sub, strings):
    """
    >>> texts = ("fail", "dolly the llama", "Hello", "Hollow", "not ok")
    >>> list(findall_each(findall, 'll', texts))
    [(), (2, 10), (2,), (2,), ()]
    >>> texts = ("parallellized", "illegally", "dillydallying", "hillbillies")
    >>> list(findall_each(findall_iter, 'll', texts))
    [(4, 7), (1, 6), (2, 7), (2, 6)]
    """
    return (tuple(findall(sub, string)) for string in strings)

3

목록 예 :

In [1]: x = ['ll','ok','ll']

In [2]: for idx, value in enumerate(x):
   ...:     if value == 'll':
   ...:         print idx, value       
0 ll
2 ll

'll'이 포함 된 목록의 모든 항목을 원하면 그렇게 할 수도 있습니다.

In [3]: x = ['Allowed','Hello','World','Hollow']

In [4]: for idx, value in enumerate(x):
   ...:     if 'll' in value:
   ...:         print idx, value
   ...:         
   ...:         
0 Allowed
1 Hello
3 Hollow

2
>>> for n,c in enumerate(text):
...   try:
...     if c+text[n+1] == "ll": print n
...   except: pass
...
1
10
16

1

일반적으로 프로그래밍에 익숙하지 않고 온라인 자습서를 통해 작업합니다. 이 작업도 요청 받았지만 지금까지 배운 방법 (기본적으로 문자열과 루프) 만 사용했습니다. 이것이 여기에 가치를 추가하는지 확실하지 않으며 이것이 당신이하는 방법이 아니라는 것을 알고 있지만 이것과 함께 작동합니다.

needle = input()
haystack = input()
counter = 0
n=-1
for i in range (n+1,len(haystack)+1):
   for j in range(n+1,len(haystack)+1):
      n=-1
      if needle != haystack[i:j]:
         n = n+1
         continue
      if needle == haystack[i:j]:
         counter = counter + 1
print (counter)

1

이 버전은 문자열 길이가 선형이어야하며 시퀀스가 ​​너무 반복적이지 않은 한 괜찮습니다 (이 경우 재귀를 while 루프로 바꿀 수 있음).

def find_all(st, substr, start_pos=0, accum=[]):
    ix = st.find(substr, start_pos)
    if ix == -1:
        return accum
    return find_all(st, substr, start_pos=ix + 1, accum=accum + [ix])

bstpierre의 list comprehension은 짧은 시퀀스에 대한 좋은 솔루션이지만 2 차 복잡도를 가지고 있고 내가 사용하던 긴 텍스트로 완성되지 않은 것 같습니다.

findall_lc = lambda txt, substr: [n for n in xrange(len(txt))
                                   if txt.find(substr, n) == n]

사소하지 않은 길이의 임의 문자열에 대해 두 함수는 동일한 결과를 제공합니다.

import random, string; random.seed(0)
s = ''.join([random.choice(string.ascii_lowercase) for _ in range(100000)])

>>> find_all(s, 'th') == findall_lc(s, 'th')
True
>>> findall_lc(s, 'th')[:4]
[564, 818, 1872, 2470]

하지만 2 차 버전은 약 300 배 더 느립니다.

%timeit find_all(s, 'th')
1000 loops, best of 3: 282 µs per loop

%timeit findall_lc(s, 'th')    
10 loops, best of 3: 92.3 ms per loop

0
#!/usr/local/bin python3
#-*- coding: utf-8 -*-

main_string = input()
sub_string = input()

count = counter = 0

for i in range(len(main_string)):
    if main_string[i] == sub_string[0]:
        k = i + 1
        for j in range(1, len(sub_string)):
            if k != len(main_string) and main_string[k] == sub_string[j]:
                count += 1
                k += 1
        if count == (len(sub_string) - 1):
            counter += 1
        count = 0

print(counter) 

이 프로그램은 정규식을 사용하지 않고 겹친 경우에도 모든 하위 문자열의 수를 계산합니다. 그러나 이것은 순진한 구현이며 최악의 경우 더 나은 결과를 얻으려면 Suffix Tree, KMP 및 기타 문자열 일치 데이터 구조 및 알고리즘을 사용하는 것이 좋습니다.


0

다음은 여러 발생을 찾는 기능입니다. 여기의 다른 솔루션과 달리 다음과 같이 슬라이싱을위한 선택적 시작 및 종료 매개 변수를 지원합니다 str.index.

def all_substring_indexes(string, substring, start=0, end=None):
    result = []
    new_start = start
    while True:
        try:
            index = string.index(substring, new_start, end)
        except ValueError:
            return result
        else:
            result.append(index)
            new_start = index + len(substring)

0

부분 문자열이 발생하는 인덱스 목록을 반환하는 간단한 반복 코드입니다.

        def allindices(string, sub):
           l=[]
           i = string.find(sub)
           while i >= 0:
              l.append(i)
              i = string.find(sub, i + 1)
           return l

0

상대 위치를 얻기 위해 분할 한 다음 목록의 연속 숫자를 합하고 동시에 (문자열 길이 * 발생 순서)를 추가하여 원하는 문자열 인덱스를 얻을 수 있습니다.

>>> key = 'll'
>>> text = "Allowed Hello Hollow"
>>> x = [len(i) for i in text.split(key)[:-1]]
>>> [sum(x[:i+1]) + i*len(key) for i in range(len(x))]
[1, 10, 16]
>>> 

0

아마도 파이썬 적이지는 않지만 좀 더 자명하다. 원래 문자열에서 본 단어의 위치를 ​​반환합니다.

def retrieve_occurences(sequence, word, result, base_counter):
     indx = sequence.find(word)
     if indx == -1:
         return result
     result.append(indx + base_counter)
     base_counter += indx + len(word)
     return retrieve_occurences(sequence[indx + len(word):], word, result, base_counter)

0

텍스트 길이를 테스트 할 필요가 없다고 생각합니다. 찾을 것이 남지 않을 때까지 계속 찾으십시오. 이렇게 :

    >>> text = 'Allowed Hello Hollow'
    >>> place = 0
    >>> while text.find('ll', place) != -1:
            print('ll found at', text.find('ll', place))
            place = text.find('ll', place) + 2


    ll found at 1
    ll found at 10
    ll found at 16

0

다음과 같이 조건부 목록 이해로 할 수도 있습니다.

string1= "Allowed Hello Hollow"
string2= "ll"
print [num for num in xrange(len(string1)-len(string2)+1) if string1[num:num+len(string2)]==string2]
# [1, 10, 16]

0

얼마 전에이 아이디어를 무작위로 얻었습니다. 문자열 스 플라이 싱 및 문자열 검색과 함께 While 루프를 사용하면 문자열이 겹치는 경우에도 작동 할 수 있습니다.

findin = "algorithm alma mater alison alternation alpines"
search = "al"
inx = 0
num_str = 0

while True:
    inx = findin.find(search)
    if inx == -1: #breaks before adding 1 to number of string
        break
    inx = inx + 1
    findin = findin[inx:] #to splice the 'unsearched' part of the string
    num_str = num_str + 1 #counts no. of string

if num_str != 0:
    print("There are ",num_str," ",search," in your string.")
else:
    print("There are no ",search," in your string.")

저는 Python 프로그래밍 (실제로 모든 언어의 프로그래밍)의 아마추어이고 다른 문제가있을 수 있는지 잘 모르겠지만 제대로 작동하는 것 같습니다.

필요한 경우 lower () 어딘가에서도 사용할 수 있다고 생각합니다.


0

다음 함수는 각 발생이 발견 된 위치를 알리면서 다른 문자열 내부에서 문자열의 모든 발생을 찾습니다.

아래 표의 테스트 케이스를 사용하여 함수를 호출 할 수 있습니다. 단어, 공백 및 숫자를 모두 섞어 시도 할 수 있습니다.

이 기능은 캐릭터가 겹치는 경우 잘 작동합니다.

|         theString          | aString |
| -------------------------- | ------- |
| "661444444423666455678966" |  "55"   |
| "661444444423666455678966" |  "44"   |
| "6123666455678966"         |  "666"  |
| "66123666455678966"        |  "66"   |

Calling examples:
1. print("Number of occurrences: ", find_all("123666455556785555966", "5555"))
   
   output:
           Found in position:  7
           Found in position:  14
           Number of occurrences:  2
   
2. print("Number of occorrences: ", find_all("Allowed Hello Hollow", "ll "))

   output:
          Found in position:  1
          Found in position:  10
          Found in position:  16
          Number of occurrences:  3

3. print("Number of occorrences: ", find_all("Aaa bbbcd$#@@abWebbrbbbbrr 123", "bbb"))

   output:
         Found in position:  4
         Found in position:  21
         Number of occurrences:  2
         

def find_all(theString, aString):
    count = 0
    i = len(aString)
    x = 0

    while x < len(theString) - (i-1): 
        if theString[x:x+i] == aString:        
            print("Found in position: ", x)
            x=x+i
            count=count+1
        else:
            x=x+1
    return count
당사 사이트를 사용함과 동시에 당사의 쿠키 정책개인정보 보호정책을 읽고 이해하였음을 인정하는 것으로 간주합니다.
Licensed under cc by-sa 3.0 with attribution required.