Tried context compression with faster models today - works surprisingly well and cuts costs significantly. The performance gains are noticeable, especially when dealing with large context windows. Makes you wonder why this isn't the default approach for most pipelines.
