# Byte\_extract / byte\_test string limits

**URL:** https://forum.suricata.io/t/byte-extract-byte-test-string-limits/4511
**Category:** Rules
**Created:** [March 1, 2024, 9:13pm UTC](https://forum.suricata.io/t/byte-extract-byte-test-string-limits/4511 "2024-03-01T21:13:16Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![bmurphy](https://avatars.discourse-cdn.com/v4/letter/b/f19dbf/32.png) [@bmurphy](https://forum.suricata.io/u/bmurphy)
#### Post date: [March 1, 2024, 9:13pm UTC](https://forum.suricata.io/t/byte-extract-byte-test-string-limits/4511/1 "2024-03-01T21:13:16Z")

</div>

I was attempting to write a rule that uses byte\_extract and byte\_test to validate that a 32 byte string extracted the http.uri buffer was found again in the http.cookie buffer.

When I attempted this following logic

```auto
http.uri; content:"&foo="; byte_extract:32,0,TESTrelative,string; http.cookie; content:"foo="; startswith; byte_test:32,=,TEST,0,relative,string;

```

I was presented with the following error:

```auto
<Error> - [ERRCODE: SC_ERR_INVALID_SIGNATURE(39)] - byte_extract can't process more than 20 bytes in "string" extraction

```

When I looked at the code i found the following constants defined within detect-byte\_extract.c

> <https://github.com/OISF/suricata/blob/6d0e11e76c8e02deada688b767523512b70e51ec/src/detect-byte-extract.c#L71-L76>

I’m left wondering a couple things:

1. What is the reasoning/background behind these limitations?

2. Is there a better way to enforce that an extracted content match between multiple buffers?

---

<div class="post-metadata">

### Author: ![Jeff\_Lucovsky](https://yyz2.discourse-cdn.com/flex030/user_avatar/forum.suricata.io/jeff_lucovsky/32/11_2.png) [@Jeff\_Lucovsky](https://forum.suricata.io/u/Jeff_Lucovsky)
#### Post date: [March 2, 2024, 2:30pm UTC](https://forum.suricata.io/t/byte-extract-byte-test-string-limits/4511/2 "2024-03-02T14:30:08Z")

</div>

Regarding #1, I’m guessing it’s a limitation carried over from snort. Snort 3 continues to limit the `count` parameter to values [1-10] (when `string` is used) and [1-4] otherwise.

The limitation exists because `byte_extract` has always been used to extract numeric quantities (hence the 1-4 and other value limitations) instead of being a general purpose “byte extraction mechanism”. The numeric values are normally used with `byte_jump`, et. al.

We have a mechanism that _may_ help – flowvars – but I don’t think that will allow the comparison logic to work the way you’d like.

That said, we could make a change to `byte_extract` to extract “byte buffers” with restrictions to prevent the value from being used in places where a numeric value is expected.

Thoughts @zoomequipd ?

---

<div class="post-metadata">

### Author: ![bmurphy](https://avatars.discourse-cdn.com/v4/letter/b/f19dbf/32.png) [@bmurphy](https://forum.suricata.io/u/bmurphy)
#### Post date: [March 4, 2024, 5:48pm UTC](https://forum.suricata.io/t/byte-extract-byte-test-string-limits/4511/3 "2024-03-04T17:48:27Z")

</div>

> We have a mechanism that may help – flowvars – but I don’t think that will allow the comparison logic to work the way you’d like.

I did check into those, and I agree, they don’t quite meet the use case.

> That said, we could make a change to `byte_extract` to extract “byte buffers” with restrictions to prevent the value from being used in places where a numeric value is expected.

I’m all for it!

I found additional examples oft his scattered in the ET ruleset. One of the more common ones was used within Phishing sigs and utilized PCRE capture groups on the http.header buffer to compare the host extracted from the referer header and compares it to the host header

I’ll get a feature submitted.

---

<div class="post-metadata">

### Author: ![bmurphy](https://avatars.discourse-cdn.com/v4/letter/b/f19dbf/32.png) [@bmurphy](https://forum.suricata.io/u/bmurphy)
#### Post date: [March 5, 2024, 6:13pm UTC](https://forum.suricata.io/t/byte-extract-byte-test-string-limits/4511/4 "2024-03-05T18:13:31Z")

</div>

> **[Feature #6831: support extraction of bytes of non-numeric values - Suricata -...](https://redmine.openinfosecfoundation.org/issues/6831)**
>
> Redmine

Wasn’t too sure how to title this request, but ticket here.
